Mathematical Reasoning on GSM8K (Acc, TPF, AUP)
96AccuracyOrthrus-Qwen3-8B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Orthrus-Qwen3-8BParams=8.7B2026.05 | 96 | — | — | |
| Initial large teacherModel=Qwen2.5-7B-Instruct2026.05 | 94.9 | — | — | |
| COSEBackbone=Qwen3-4B2026.05 | 94.4 | — | — | |
| AZRBackbone=Qwen3-4B2026.05 | 93.2 | — | — | |
| COSEBackbone=Qwen3-0.6B2026.05 | 93 | — | — | |
| MAEBackbone=Qwen3-4B2026.05 | 92.8 | — | — | |
| RL (BF16)Model=Qwen3.5-9B, Bitwidth=BF162026.05 | 92.57 | — | — | |
| CGSModel pair=Qwen2.5-1.5B → Qwen2.5-32B, k=32, Chunk length (L)=202026.06 | 92.5 | — | — | |
| R-ZeroBackbone=Qwen3-4B2026.05 | 92.4 | — | — | |
| PRM guided searchModel pair=Qwen2.5-1.5B → Qwen2.5-32B, k=32, Guidance=Reasoning steps2026.06 | 92.3 | — | — | |
| AISModel=Qwen3.5-9B, Bitwidth=FP82026.05 | 91.74 | — | — | |
| SDAR-Qwen3-8BParams=8B2026.05 | 91.7 | — | — | |
| SDAR-Qwen3-30B-A3BParams=30B2026.05 | 91.4 | — | — | |
| FlashRL (TIS)Model=Qwen3.5-9B, Bitwidth=FP82026.05 | 90.8 | — | — | |
| LGSModel pair=Qwen2.5-1.5B → Qwen2.5-32B, k=32, Chunk length (L)=202026.06 | 90.4 | — | — | |
| LightningRL-8B-b32Model Category=Block-wise dLLMs, Evaluation Protocol=Zero-shot2026.03 | 90.3 | 5.58 | 492.4 | |
| RL (FP8 Rollout)Model=Qwen3.5-9B, Bitwidth=FP82026.05 | 90.22 | — | — | |
| AISModel=Qwen3-8B, Bitwidth=FP82026.05 | 89.99 | — | — | |
| RL (BF16)Model=Qwen3-8B, Bitwidth=BF162026.05 | 89.84 | — | — | |
| CGSModel pair=Qwen2.5-1.5B → Qwen2.5-32B, k=16, Chunk length (L)=202026.06 | 89.3 | — | — | |
| PRM guided searchModel pair=Qwen2.5-1.5B → Qwen2.5-32B, k=16, Guidance=Reasoning steps2026.06 | 89.2 | — | — | |
| BoostLoRA# Additional Params=122026.04 | 89.1 | — | — | |
| COSEBackbone=Llama-3.2-3B-Instruct2026.05 | 89 | — | — | |
| SDAR-8B-b32Model Category=Block-wise dLLMs, Evaluation Protocol=Zero-shot2026.03 | 88.9 | 2.85 | 252.5 | |
| FlashRL (TIS)Model=Qwen3-8B, Bitwidth=FP82026.05 | 88.63 | — | — | |
| Initial small teacherModel=Qwen2.5-3B-Instruct2026.05 | 88.5 | — | — | |
| RL (FP8 Rollout)Model=Qwen3-8B, Bitwidth=FP82026.05 | 87.65 | — | — | |
| LGSModel pair=Qwen2.5-1.5B → Qwen2.5-32B, k=16, Chunk length (L)=202026.06 | 87.6 | — | — | |
| COSEBackbone=Qwen2.5-3B-Instruct2026.05 | 87.33 | — | — | |
| TinyLoRA# Additional Params=8,0642026.04 | 87.2 | — | — | |
| Full FT# Additional Params=3.09B2026.04 | 87 | — | — | |
| MAEBackbone=Qwen2.5-3B-Instruct2026.05 | 86.8 | — | — | |
| TinyLoRA# Additional Params=129,0242026.04 | 86.7 | — | — | |
| AZRBackbone=Qwen2.5-3B-Instruct2026.05 | 86.4 | — | — | |
| R-ZeroBackbone=Qwen2.5-3B-Instruct2026.05 | 86 | — | — | |
| PRM guided searchModel pair=Qwen2.5-1.5B → Qwen2.5-32B, k=8, Guidance=Reasoning steps2026.06 | 86 | — | — | |
| CGSModel pair=Qwen2.5-1.5B → Qwen2.5-32B, k=8, Chunk length (L)=202026.06 | 85.7 | — | — | |
| MAEBackbone=Qwen3-0.6B2026.05 | 85.6 | — | — | |
| TinyLoRA# Additional Params=2522026.04 | 85.4 | — | — | |
| LGSModel pair=Qwen2.5-1.5B → Qwen2.5-32B, k=8, Chunk length (L)=202026.06 | 85.4 | — | — | |
| BaseBackbone=Qwen2.5-3B-Instruct2026.05 | 85.2 | — | — | |
| BaseBackbone=Qwen3-4B2026.05 | 85 | — | — | |
| R-ZeroBackbone=Qwen3-0.6B2026.05 | 84.8 | — | — | |
| AZRBackbone=Qwen3-0.6B2026.05 | 84.2 | — | — | |
| DreamModel Category=dLLMs, Evaluation Protocol=Zero-shot2026.03 | 83.9 | 1 | 83.9 | |
| CGSModel pair=Llama-3.2-1B → Llama-3.1-70B, k=32, Chunk length (L)=202026.06 | 83.9 | — | — | |
| Fast-dLLM-v2Params=7B2026.05 | 83.7 | — | — | |
| LLaDA-1.5Params=7B2026.05 | 82.4 | — | — | |
| dParallel-DreamModel Category=dLLMs, Evaluation Protocol=Zero-shot2026.03 | 82.1 | 3.02 | 245.7 | |
| d3LLM-DreamModel Category=dLLMs, Evaluation Protocol=Zero-shot2026.03 | 81.4 | 4.94 | 391.3 | |
| LGSModel pair=Llama-3.2-1B → Llama-3.1-70B, k=32, Chunk length (L)=202026.06 | 81.2 | — | — | |
| BaseBackbone=Qwen3-0.6B2026.05 | 81 | — | — | |
| TinyLoRA# Additional Params=162026.04 | 80.9 | — | — | |
| Dream-7BParams=7B2026.05 | 79.3 | — | — | |
| CGSModel pair=Llama-3.2-1B → Llama-3.1-70B, k=16, Chunk length (L)=202026.06 | 79.3 | — | — | |
| Fast-dLLM-DreamModel Category=dLLMs, Evaluation Protocol=Zero-shot2026.03 | 79 | 1.44 | 116.5 | |
| MSA-PTTraining Strategy=from-scratch sparse pretraining2026.06 | 77.7 | — | — | |
| Fast-dLLM-v2Model Category=Block-wise dLLMs, Evaluation Protocol=Zero-shot2026.03 | 77.5 | 2.21 | 156 | |
| PRM guided searchModel pair=Llama-3.2-1B → Llama-3.1-70B, k=32, Guidance=Reasoning steps2026.06 | 77.3 | — | — | |
| LGSModel pair=Llama-3.2-1B → Llama-3.1-70B, k=16, Chunk length (L)=202026.06 | 77.3 | — | — | |
| EAGLE-3 (LLaMA-3.1)Model Category=AR Models, Evaluation Protocol=Zero-shot2026.03 | 76.6 | 5.12 | 319 | |
| FullTraining Strategy=Full-Attention baseline2026.06 | 76.2 | — | — | |
| Base modelmode=zero-shot, # Additional Params=02026.04 | 76 | — | — | |
| MAEBackbone=Llama-3.2-3B-Instruct2026.05 | 75.2 | — | — | |
| Fast-dLLM-LLaDAModel Category=dLLMs, Evaluation Protocol=Zero-shot2026.03 | 74.7 | 2.77 | 205.8 | |
| Qwen-2.5-7B-itModel Category=AR Models, Evaluation Protocol=Zero-shot2026.03 | 74.1 | 1 | 74.1 | |
| R-ZeroBackbone=Llama-3.2-3B-Instruct2026.05 | 73.8 | — | — | |
| MSA-CPTTraining Strategy=sparse continued pretraining2026.06 | 73.7 | — | — | |
| D2F-LLaDAModel Category=dLLMs, Evaluation Protocol=Zero-shot2026.03 | 73.2 | 2.88 | 209.7 | |
| d3LLM-LLaDAModel Category=dLLMs, Evaluation Protocol=Zero-shot2026.03 | 73.1 | 9.11 | 637.7 | |
| LLaDAModel Category=dLLMs, Evaluation Protocol=Zero-shot2026.03 | 72.6 | 1 | 72.6 | |
| dParallel-LLaDAModel Category=dLLMs, Evaluation Protocol=Zero-shot2026.03 | 72.6 | 5.14 | 358.1 | |
| AZRBackbone=Llama-3.2-3B-Instruct2026.05 | 72.4 | — | — | |
| PRM guided searchModel pair=Llama-3.2-1B → Llama-3.1-70B, k=16, Guidance=Reasoning steps2026.06 | 71.8 | — | — | |
| CGSModel pair=Llama-3.2-1B → Llama-3.1-70B, k=8, Chunk length (L)=202026.06 | 71.3 | — | — | |
| LGSModel pair=Llama-3.2-1B → Llama-3.1-70B, k=8, Chunk length (L)=202026.06 | 69.8 | — | — | |
| PRM guided searchModel pair=Llama-3.2-1B → Llama-3.1-70B, k=8, Guidance=Reasoning steps2026.06 | 65.4 | — | — | |
| BaseBackbone=Llama-3.2-3B-Instruct2026.05 | 63 | — | — | |
| CLPDStudent Model=Qwen-0.5B, Data Ordering=expert CoT2026.05 | 48.5 | — | — | |
| CLPDStudent Model=Qwen-0.5B, Data Ordering=student-loss2026.05 | 48.2 | — | — | |
| CLPDStudent Model=Llama-1B, Data Ordering=expert CoT2026.05 | 46.5 | — | — | |
| CLPDStudent Model=Llama-1B, Data Ordering=student-loss2026.05 | 46.3 | — | — | |
| Progressive distillation onlyStudent Model=Qwen-0.5B, Teacher=small & large teacher2026.05 | 46 | — | — | |
| Curriculum learning onlyStudent Model=Qwen-0.5B, Teacher=small teacher, Data Ordering=expert CoT2026.05 | 45.7 | — | — | |
| Curriculum learning onlyStudent Model=Qwen-0.5B, Teacher=small teacher, Data Ordering=student-loss2026.05 | 45.2 | — | — | |
| Curriculum learning onlyStudent Model=Qwen-0.5B, Teacher=large teacher, Data Ordering=expert CoT2026.05 | 45.1 | — | — | |
| Curriculum learning onlyStudent Model=Qwen-0.5B, Teacher=large teacher, Data Ordering=student-loss2026.05 | 45.1 | — | — | |
| Standard distillationStudent Model=Qwen-0.5B, Teacher=small teacher2026.05 | 44.9 | — | — | |
| Progressive distillation onlyStudent Model=Llama-1B, Teacher=small & large teacher2026.05 | 44.9 | — | — | |
| Standard distillationStudent Model=Qwen-0.5B, Teacher=large teacher2026.05 | 44.6 | — | — | |
| Curriculum learning onlyStudent Model=Llama-1B, Teacher=small teacher, Data Ordering=expert CoT2026.05 | 44.6 | — | — | |
| Curriculum learning onlyStudent Model=Llama-1B, Teacher=large teacher, Data Ordering=student-loss2026.05 | 44.4 | — | — | |
| Curriculum learning onlyStudent Model=Llama-1B, Teacher=large teacher, Data Ordering=expert CoT2026.05 | 44 | — | — | |
| Curriculum learning onlyStudent Model=Llama-1B, Teacher=small teacher, Data Ordering=student-loss2026.05 | 43.8 | — | — | |
| Standard distillationStudent Model=Llama-1B, Teacher=small teacher2026.05 | 42.7 | — | — | |
| Standard distillationStudent Model=Llama-1B, Teacher=large teacher2026.05 | 42.4 | — | — | |
| Initial studentStudent Model=Qwen-0.5B2026.05 | 41.4 | — | — | |
| Initial studentStudent Model=Llama-1B2026.05 | 40.6 | — | — |