Mathematical Reasoning on AIME 24 (pass@2, pass@4, pass@8, pass@16)
27.03Pass@2Qwen3-235B-A22B-Instr.
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Qwen3-235B-A22B-Instr.Parameter Scale=235B, Training Strategy=Commercial Baseline2026.03 | 27.03 | 28.32 | 30 | 33.33 | |
| DeepSeek-R1-671B-0528Parameter Scale=671B, Training Strategy=Commercial Baseline2026.03 | 21.03 | 28.07 | 32.14 | 33.33 | |
| Qwen2.5-7B-ERPOParameter Scale=7B, Training Strategy=ERPO2026.03 | 17.89 | 22.45 | 27.38 | 33.33 | |
| Qwen2.5-7B-GRPOParameter Scale=7B, Training Strategy=GRPO2026.03 | 16.28 | 21.49 | 27.09 | 33.33 | |
| Qwen2.5-3B-ERPOParameter Scale=3B, Training Strategy=ERPO2026.03 | 11.22 | 16.43 | 23.4 | 33.33 | |
| Qwen2.5-3B-GRPOParameter Scale=3B, Training Strategy=GRPO2026.03 | 8.58 | 12.81 | 18.44 | 26.67 | |
| Qwen2.5-1.5B-ERPOParameter Scale=1.5B, Training Strategy=ERPO2026.03 | 6.81 | 11.4 | 17.08 | 23.33 | |
| Qwen2.5-1.5B-GRPOParameter Scale=1.5B, Training Strategy=GRPO2026.03 | 6.19 | 10.06 | 15.56 | 23.33 | |
| Qwen2.5-7B-BaseParameter Scale=7B, Training Strategy=Base2026.03 | 6.19 | 10.06 | 15.56 | 23.33 | |
| Qwen2.5-3B-BaseParameter Scale=3B, Training Strategy=Base2026.03 | 3.92 | 6.95 | 11.21 | 16.67 | |
| Qwen2.5-7B-SFTParameter Scale=7B, Training Strategy=SFT2026.03 | 2.86 | 5.5 | 10.11 | 16.67 | |
| Qwen2.5-1.5B-SFTParameter Scale=1.5B, Training Strategy=SFT2026.03 | 1.67 | 3.33 | 6.67 | 13.33 | |
| Qwen2.5-3B-SFTParameter Scale=3B, Training Strategy=SFT2026.03 | 1.64 | 3.17 | 5.89 | 10 | |
| Qwen2.5-1.5B-BaseParameter Scale=1.5B, Training Strategy=Base2026.03 | 0.42 | 0.83 | 1.67 | 3.33 |