Mathematical Reasoning on AIME’25, AIME’24, AMC’23, and MATH500 (test)
92.3Acc@4Qwen-3-30B-Thinking
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen-3-30B-ThinkingParams=30B, Role=Teacher performance2026.01 | 92.3 | |
| Qwen-3-235B-ThinkingParams=235B, Role=Teacher performance2026.01 | 91.2 | |
| Deepseek-R1Params=671B, Role=Teacher performance2026.01 | 91.1 | |
| GPT-OSS-120BParams=120B, Role=Teacher performance2026.01 | 88.3 | |
| Qwen-3-4B-ThinkingParams=4B, Role=Teacher performance2026.01 | 87.3 | |
| QwQ-32BParams=32B, Role=Teacher performance2026.01 | 85.2 | |
| GPT-OSS-20BParams=20B, Role=Teacher performance2026.01 | 83.4 | |
| Qwen-3-8BParams=8B, Role=Teacher performance2026.01 | 82.5 | |
| Nemotron-SuperParams=49B, Role=Teacher performance2026.01 | 82.3 | |
| Qwen-3-14BTeacher Model=QwQ-32B, Teacher Params=32B, Post-training distillation=true2026.01 | 77.4 | |
| Qwen-3-14BTeacher Model=Qwen-3-30B-Thinking, Teacher Params=30B, Post-training distillation=true2026.01 | 77.2 | |
| Qwen-3-14BTeacher Model=Deepseek-R1, Teacher Params=671B, Post-training distillation=true2026.01 | 77.1 | |
| Qwen-3-14BTeacher Model=Qwen-3-4B-Thinking, Teacher Params=4B, Post-training distillation=true2026.01 | 76.8 | |
| Qwen-3-14BTeacher Model=Qwen-3-8B, Teacher Params=8B, Post-training distillation=true2026.01 | 74.6 | |
| Phi-4-Reasoning-PlusParams=14B, Role=Teacher performance2026.01 | 72.7 | |
| Qwen-3-14BTeacher Model=Nemotron-Super, Teacher Params=49B, Post-training distillation=true2026.01 | 72.2 | |
| Qwen-3-14BTeacher Model=Qwen-3-235B-Thinking, Teacher Params=235B, Post-training distillation=true2026.01 | 71.8 | |
| Magistral-SmallParams=24B, Role=Teacher performance2026.01 | 71 | |
| Qwen-3-14BTeacher Model=GPT-OSS-20B, Teacher Params=20B, Post-training distillation=true2026.01 | 69.5 | |
| Qwen-3-14BTeacher Model=Magistral-Small, Teacher Params=24B, Post-training distillation=true2026.01 | 68.8 | |
| Qwen-3-14BTeacher Model=GPT-OSS-120B, Teacher Params=120B, Post-training distillation=true2026.01 | 66.7 | |
| Qwen-3-4BTeacher Model=Qwen-3-4B-Thinking, Teacher Params=4B, Post-training distillation=true2026.01 | 61.9 | |
| Qwen-3-4BTeacher Model=QwQ-32B, Teacher Params=32B, Post-training distillation=true2026.01 | 61.2 | |
| Qwen-3-4BTeacher Model=Qwen-3-8B, Teacher Params=8B, Post-training distillation=true2026.01 | 61.2 | |
| Qwen-3-4BTeacher Model=Qwen-3-30B-Thinking, Teacher Params=30B, Post-training distillation=true2026.01 | 58.8 | |
| Qwen-3-4BTeacher Model=Nemotron-Super, Teacher Params=49B, Post-training distillation=true2026.01 | 56.4 | |
| Qwen-3-4BTeacher Model=Deepseek-R1, Teacher Params=671B, Post-training distillation=true2026.01 | 55.8 | |
| Qwen-3-14BTeacher Model=Phi-4-Reasoning-Plus, Teacher Params=14B, Post-training distillation=true2026.01 | 54.1 | |
| Qwen-3-4BTeacher Model=Qwen-3-235B-Thinking, Teacher Params=235B, Post-training distillation=true2026.01 | 53.4 | |
| Qwen-3-4BTeacher Model=Magistral-Small, Teacher Params=24B, Post-training distillation=true2026.01 | 52.2 | |
| Qwen-2.5-7BTeacher Model=QwQ-32B, Teacher Params=32B, Post-training distillation=true2026.01 | 52 | |
| Qwen-2.5-7BTeacher Model=Qwen-3-8B, Teacher Params=8B, Post-training distillation=true2026.01 | 52 | |
| Qwen-2.5-7BTeacher Model=Qwen-3-4B-Thinking, Teacher Params=4B, Post-training distillation=true2026.01 | 51.8 | |
| Qwen-2.5-7BTeacher Model=Qwen-3-30B-Thinking, Teacher Params=30B, Post-training distillation=true2026.01 | 50 | |
| Qwen-3-4BTeacher Model=GPT-OSS-20B, Teacher Params=20B, Post-training distillation=true2026.01 | 48.4 | |
| Qwen-2.5-7BTeacher Model=Nemotron-Super, Teacher Params=49B, Post-training distillation=true2026.01 | 48.3 | |
| Qwen-3-4BTeacher Model=GPT-OSS-120B, Teacher Params=120B, Post-training distillation=true2026.01 | 47.9 | |
| Qwen-2.5-7BTeacher Model=Magistral-Small, Teacher Params=24B, Post-training distillation=true2026.01 | 47.6 | |
| Qwen-2.5-7BTeacher Model=Deepseek-R1, Teacher Params=671B, Post-training distillation=true2026.01 | 47.3 | |
| Qwen-2.5-7BTeacher Model=Qwen-3-235B-Thinking, Teacher Params=235B, Post-training distillation=true2026.01 | 45 | |
| Qwen-2.5-7BTeacher Model=GPT-OSS-20B, Teacher Params=20B, Post-training distillation=true2026.01 | 42.7 | |
| Qwen-2.5-7BTeacher Model=GPT-OSS-120B, Teacher Params=120B, Post-training distillation=true2026.01 | 40.7 | |
| Qwen-3-4BTeacher Model=Phi-4-Reasoning-Plus, Teacher Params=14B, Post-training distillation=true2026.01 | 40.2 | |
| Qwen-2.5-7BTeacher Model=Phi-4-Reasoning-Plus, Teacher Params=14B, Post-training distillation=true2026.01 | 35.2 | |
| Qwen-2.5-3BTeacher Model=Qwen-3-8B, Teacher Params=8B, Post-training distillation=true2026.01 | 34.2 | |
| Qwen-2.5-3BTeacher Model=Qwen-3-4B-Thinking, Teacher Params=4B, Post-training distillation=true2026.01 | 33.3 | |
| Qwen-2.5-3BTeacher Model=Nemotron-Super, Teacher Params=49B, Post-training distillation=true2026.01 | 33 | |
| Qwen-2.5-3BTeacher Model=QwQ-32B, Teacher Params=32B, Post-training distillation=true2026.01 | 33 | |
| Qwen-2.5-3BTeacher Model=Qwen-3-30B-Thinking, Teacher Params=30B, Post-training distillation=true2026.01 | 31.2 | |
| Qwen-2.5-3BTeacher Model=Magistral-Small, Teacher Params=24B, Post-training distillation=true2026.01 | 30.6 | |
| Qwen-2.5-3BTeacher Model=Deepseek-R1, Teacher Params=671B, Post-training distillation=true2026.01 | 29.6 | |
| LLaMA-3.1-8BTeacher Model=Qwen-3-4B-Thinking, Teacher Params=4B, Post-training distillation=true2026.01 | 28.2 | |
| LLaMA-3.1-8BTeacher Model=Deepseek-R1, Teacher Params=671B, Post-training distillation=true2026.01 | 28.1 | |
| LLaMA-3.1-8BTeacher Model=QwQ-32B, Teacher Params=32B, Post-training distillation=true2026.01 | 27.1 | |
| LLaMA-3.1-8BTeacher Model=Qwen-3-30B-Thinking, Teacher Params=30B, Post-training distillation=true2026.01 | 26.7 | |
| LLaMA-3.1-8BTeacher Model=Qwen-3-8B, Teacher Params=8B, Post-training distillation=true2026.01 | 26.5 | |
| Qwen-2.5-3BTeacher Model=Qwen-3-235B-Thinking, Teacher Params=235B, Post-training distillation=true2026.01 | 26.4 | |
| Qwen-2.5-3BTeacher Model=GPT-OSS-20B, Teacher Params=20B, Post-training distillation=true2026.01 | 24.4 | |
| LLaMA-3.1-8BTeacher Model=Nemotron-Super, Teacher Params=49B, Post-training distillation=true2026.01 | 23.7 | |
| Qwen-2.5-3BTeacher Model=GPT-OSS-120B, Teacher Params=120B, Post-training distillation=true2026.01 | 22.9 | |
| LLaMA-3.1-8BTeacher Model=Magistral-Small, Teacher Params=24B, Post-training distillation=true2026.01 | 22.8 | |
| LLaMA-3.1-8BTeacher Model=Qwen-3-235B-Thinking, Teacher Params=235B, Post-training distillation=true2026.01 | 22 | |
| Qwen-2.5-3BTeacher Model=Phi-4-Reasoning-Plus, Teacher Params=14B, Post-training distillation=true2026.01 | 18.2 | |
| LLaMA-3.1-8BTeacher Model=GPT-OSS-20B, Teacher Params=20B, Post-training distillation=true2026.01 | 17.9 | |
| LLaMA-3.1-8BTeacher Model=GPT-OSS-120B, Teacher Params=120B, Post-training distillation=true2026.01 | 15.2 | |
| LLaMA-3.1-8BTeacher Model=Phi-4-Reasoning-Plus, Teacher Params=14B, Post-training distillation=true2026.01 | 14.5 |