Mathematical Reasoning on AIME2024, AMC, MATH-500, Minerva, and Olympiad
33.5AIME 2024 ScoreSFT-then-RL
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| SFT-then-RLBackbone=Qwen2.5-Math-7B, Guidance=RL Incorporating External Expert Guidance2025.12 | 33.5 | 62.3 | 86.6 | 41.2 | 47.6 | 54.3 | |
| Oat-ZeroBackbone=Qwen2.5-Math-7B, Guidance=Pure RL without External Expert Guidance2025.12 | 33.4 | 61.2 | 78 | 34.6 | 43.4 | 50.1 | |
| LUFFYBackbone=Qwen2.5-Math-7B, Guidance=RL Incorporating External Expert Guidance2025.12 | 29.4 | 65.5 | 88.4 | 38.2 | 56 | 55.5 | |
| TRAPOBackbone=Qwen2.5-Math-7B2025.12 | 28.3 | 66.2 | 89.2 | 41.5 | 57.6 | 56.6 | |
| ReLIFTBackbone=Qwen2.5-Math-7B, Guidance=RL Incorporating External Expert Guidance2025.12 | 28.2 | 64.9 | 87.4 | 33.8 | 52.5 | 53.4 | |
| SFTBackbone=Qwen2.5-Math-7B, Guidance=RL Incorporating External Expert Guidance2025.12 | 27.7 | 56 | 84.8 | 38.2 | 44.7 | 50.3 | |
| SimpleRL-ZeroBackbone=Qwen2.5-Math-7B, Guidance=Pure RL without External Expert Guidance2025.12 | 27 | 54.9 | 76 | 25 | 34.7 | 43.5 | |
| GRPOBackbone=Qwen2.5-Math-7B, Guidance=Pure RL without External Expert Guidance2025.12 | 24 | 59 | 84 | 39.3 | 45.8 | 50.4 | |
| PRIME-ZeroBackbone=Qwen2.5-Math-7B, Guidance=Pure RL without External Expert Guidance2025.12 | 17 | 54 | 81.4 | 39 | 40.3 | 46.3 | |
| OpenReasoner-ZeroBackbone=Qwen2.5-Math-7B, Guidance=Pure RL without External Expert Guidance2025.12 | 16.5 | 52.1 | 82.4 | 33.1 | 47.1 | 46.2 | |
| Qwen2.5-Math-7BType=Base Model2025.12 | 11.9 | 33.7 | 47 | 12.5 | 21.9 | 25.4 | |
| Qwen2.5-Math-7B-InstructType=Instruct Model2025.12 | 11.3 | 48.2 | 82.6 | 36.8 | 39.7 | 43.7 |