Mathematical Reasoning on GSM8K (Accuracy, Average)
89.9AccuracyQwen2.5-7B-Instruct
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen2.5-7B-InstructArchitecture=Qwen2.5-7B, Training Stage=AR LLM2026.06 | 89.9 | 54.5 | |
| AGDO-RLArchitecture=Dream-v0-Instruct-7B, Training Stage=RL, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 87.7 | 38.1 | |
| AGDOArchitecture=Dream-v0-Instruct-7B, Training Stage=RL, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 86.9 | 38.8 | |
| TraceRLArchitecture=Dream-v0-Instruct-7B, Training Stage=RL, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 86.3 | 36.5 | |
| Coupled RLArchitecture=Dream-v0-Instruct-7B, Training Stage=RL, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 86.1 | 34.9 | |
| blockwise SFTArchitecture=Dream-v0-Instruct-7B, Training Stage=SFT, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 86 | 34.4 | |
| AGDO-SFTArchitecture=Dream-v0-Instruct-7B, Training Stage=SFT, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 85.3 | 36 | |
| Diff-GRPOArchitecture=Dream-v0-Instruct-7B, Training Stage=RL, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 85 | 35 | |
| Llama3.1-8B-InstructArchitecture=Llama3.1-8B, Training Stage=AR LLM2026.06 | 84.5 | 42.7 | |
| SFTArchitecture=Dream-v0-Instruct-7B, Training Stage=SFT, Sampling Temperature=0.1, Decoding Strategy=Static decoding, Max Response Length=1,0242026.06 | 83.5 | 33.9 | |
| LLaDA-8B-InstructArchitecture=LLaDA-8B, Training Stage=Masked DLLM2026.06 | 81.5 | 28.5 | |
| Dream-v0-Instruct-7BArchitecture=Dream-v0-Instruct-7B, Training Stage=Masked DLLM2026.06 | 69.4 | 28.3 |