Mathematical Reasoning on Minerva (pass@2, pass@4, pass@8, pass@16)
19.5Pass@2Qwen3-235B-A22B-Instr.
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Qwen3-235B-A22B-Instr.Parameter Scale=235B, Training Strategy=Commercial Baseline2026.03 | 19.5 | 21.44 | 23.4 | 25 | |
| Qwen2.5-7B-ERPOParameter Scale=7B, Training Strategy=ERPO2026.03 | 17.92 | 23.1 | 28.29 | 32.5 | |
| Qwen2.5-7B-GRPOParameter Scale=7B, Training Strategy=GRPO2026.03 | 17.83 | 24.15 | 30.54 | 35 | |
| DeepSeek-R1-671B-0528Parameter Scale=671B, Training Strategy=Commercial Baseline2026.03 | 16.35 | 20.94 | 23.78 | 25 | |
| Qwen2.5-3B-ERPOParameter Scale=3B, Training Strategy=ERPO2026.03 | 13.25 | 17.73 | 22.57 | 27.5 | |
| Qwen2.5-3B-GRPOParameter Scale=3B, Training Strategy=GRPO2026.03 | 10.94 | 14.58 | 17.56 | 20 | |
| Qwen2.5-7B-SFTParameter Scale=7B, Training Strategy=SFT2026.03 | 8.17 | 12.31 | 16.47 | 20 | |
| Qwen2.5-7B-BaseParameter Scale=7B, Training Strategy=Base2026.03 | 7.75 | 12.46 | 17.6 | 22.5 | |
| Qwen2.5-1.5B-GRPOParameter Scale=1.5B, Training Strategy=GRPO2026.03 | 6.92 | 10.4 | 13.62 | 17.5 | |
| Qwen2.5-1.5B-ERPOParameter Scale=1.5B, Training Strategy=ERPO2026.03 | 6.21 | 8.53 | 11.92 | 17.5 | |
| Qwen2.5-3B-BaseParameter Scale=3B, Training Strategy=Base2026.03 | 5.44 | 8.62 | 12.24 | 15 | |
| Qwen2.5-3B-SFTParameter Scale=3B, Training Strategy=SFT2026.03 | 3.77 | 6.56 | 10.38 | 15 | |
| Qwen2.5-1.5B-SFTParameter Scale=1.5B, Training Strategy=SFT2026.03 | 2.48 | 3.96 | 5.66 | 7.5 | |
| Qwen2.5-1.5B-BaseParameter Scale=1.5B, Training Strategy=Base2026.03 | 0.62 | 1.25 | 2.5 | 5 |