Mathematical Reasoning on AMC 23 (Pass@2, Pass@4, Pass@8, Pass@16)
63.4Pass@2Qwen2.5-7B-ERPO
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Qwen2.5-7B-ERPOParameter Scale=7B, Training Strategy=ERPO2026.03 | 63.4 | 75.22 | 84.21 | 90 | |
| Qwen2.5-7B-GRPOParameter Scale=7B, Training Strategy=GRPO2026.03 | 60.81 | 71.62 | 79.25 | 82.5 | |
| Qwen3-235B-A22B-Instr.Parameter Scale=235B, Training Strategy=Commercial Baseline2026.03 | 54.35 | 59.66 | 63 | 65 | |
| Qwen2.5-3B-ERPOParameter Scale=3B, Training Strategy=ERPO2026.03 | 49.79 | 61.91 | 71.11 | 77.5 | |
| DeepSeek-R1-671B-0528Parameter Scale=671B, Training Strategy=Commercial Baseline2026.03 | 47.62 | 58.72 | 66.19 | 72.5 | |
| Qwen2.5-3B-GRPOParameter Scale=3B, Training Strategy=GRPO2026.03 | 46 | 59.57 | 71.73 | 80 | |
| Qwen2.5-1.5B-ERPOParameter Scale=1.5B, Training Strategy=ERPO2026.03 | 38.25 | 49.62 | 61.46 | 72.5 | |
| Qwen2.5-7B-BaseParameter Scale=7B, Training Strategy=Base2026.03 | 38.15 | 54.06 | 67.42 | 77.5 | |
| Qwen2.5-1.5B-GRPOParameter Scale=1.5B, Training Strategy=GRPO2026.03 | 36.67 | 49.09 | 62.92 | 75 | |
| Qwen2.5-7B-SFTParameter Scale=7B, Training Strategy=SFT2026.03 | 28.56 | 43.31 | 59.2 | 70 | |
| Qwen2.5-3B-BaseParameter Scale=3B, Training Strategy=Base2026.03 | 25.29 | 38.6 | 52.48 | 67.5 | |
| Qwen2.5-3B-SFTParameter Scale=3B, Training Strategy=SFT2026.03 | 18.31 | 28.64 | 40.35 | 52.5 | |
| Qwen2.5-1.5B-SFTParameter Scale=1.5B, Training Strategy=SFT2026.03 | 14.9 | 25.41 | 39.17 | 55 | |
| Qwen2.5-1.5B-BaseParameter Scale=1.5B, Training Strategy=Base2026.03 | 1.56 | 3.12 | 6.25 | 12.5 |