Mathematical Reasoning on AMC 23 (Pass@1, Pass@8)
72Pass@1Rule+CER
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Rule+CERBackbone=Qwen3-8B-Base, Training Dataset=mathematical dataset2026.03 | 72 | — | |
| CERBackbone=Qwen3-8B-Base, Training Dataset=mathematical dataset2026.03 | 70.9 | — | |
| Rule-basedBackbone=Qwen3-8B-Base, Training Dataset=mathematical dataset2026.03 | 70.2 | — | |
| General-verifierBackbone=Qwen3-8B-Base, Training Dataset=mathematical dataset2026.03 | 69.4 | — | |
| VeriFreeBackbone=Qwen3-8B-Base, Training Dataset=mathematical dataset2026.03 | 68.4 | — | |
| Rule+CERBackbone=Qwen3-4B-Base, Training Dataset=mathematical dataset2026.03 | 67.5 | — | |
| CERBackbone=Qwen3-4B-Base, Training Dataset=mathematical dataset2026.03 | 63.6 | — | |
| Exact-matchBackbone=Qwen3-8B-Base, Training Dataset=mathematical dataset2026.03 | 63.4 | — | |
| Rule-basedBackbone=Qwen3-4B-Base, Training Dataset=mathematical dataset2026.03 | 63.1 | — | |
| General-verifierBackbone=Qwen3-4B-Base, Training Dataset=mathematical dataset2026.03 | 63 | — | |
| VeriFreeBackbone=Qwen3-4B-Base, Training Dataset=mathematical dataset2026.03 | 62.7 | — | |
| MTTeacher Model=Qwen2.5-Math-7B-Instruct2026.04 | 60 | — | |
| VCRDTeacher Model=Qwen2.5-Math-7B-Instruct, Student Model=Qwen2.5-Math-1.5B2026.04 | 60 | — | |
| Exact-matchBackbone=Qwen3-4B-Base, Training Dataset=mathematical dataset2026.03 | 57.3 | — | |
| BaseBackbone=Qwen3-8B-Base, Training Dataset=mathematical dataset2026.03 | 53.1 | — | |
| GKDTeacher Model=Qwen2.5-Math-7B-Instruct, Student Model=Qwen2.5-Math-1.5B2026.04 | 50 | — | |
| KDTeacher Model=Qwen2.5-Math-7B-Instruct, Student Model=Qwen2.5-Math-1.5B2026.04 | 47.5 | — | |
| DistilLLMTeacher Model=Qwen2.5-Math-7B-Instruct, Student Model=Qwen2.5-Math-1.5B2026.04 | 47.5 | — | |
| ABKDTeacher Model=Qwen2.5-Math-7B-Instruct, Student Model=Qwen2.5-Math-1.5B2026.04 | 47.5 | — | |
| DistillLM-2Teacher Model=Qwen2.5-Math-7B-Instruct, Student Model=Qwen2.5-Math-1.5B2026.04 | 47.5 | — | |
| BHA2026.04 | 47.2 | 72.5 | |
| DAPO2026.04 | 46.9 | 77.5 | |
| BREAD2026.04 | 46.6 | 70 | |
| BaseBackbone=Qwen3-4B-Base, Training Dataset=mathematical dataset2026.03 | 40.5 | — | |
| Base2026.04 | 25 | 70 | |
| MSStudent Model=Qwen2.5-Math-1.5B2026.04 | 22.5 | — | |
| GRPO (Ts : 0.9 -> 1.5)Sampling temperature (Ts)=0.9 -> 1.5, Maximum response length=6k tokens, Training algorithm=GRPO2026.02 | 0.55 | 0.775 | |
| TAMPOMaximum response length=6k tokens, Training algorithm=TAMPO2026.02 | 0.55 | 0.825 | |
| GRPO (Ts : 1.5)Sampling temperature (Ts)=1.5, Maximum response length=6k tokens, Training algorithm=GRPO2026.02 | 0.525 | 0.775 | |
| GRPO (Ts : 0.9)Sampling temperature (Ts)=0.9, Maximum response length=6k tokens, Training algorithm=GRPO2026.02 | 0.5 | 0.8 | |
| GRPO (Ts : 1.2)Sampling temperature (Ts)=1.2, Maximum response length=6k tokens, Training algorithm=GRPO2026.02 | 0.5 | 0.775 | |
| DS-Qwen-1.5BMaximum response length=6k tokens2026.02 | 0.45 | 0.725 |