Mathematical Reasoning on AIME 2024 (AVG@32, Pass@32)
79.56AVG@32OLMo-3.1-32B-Think
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| OLMo-3.1-32B-ThinkModel Size=32B, Training Recipe Category=The SOTA Training Recipes, Type=RL2026.07 | 79.56 | 93.33 | |
| Projected (Ours)Model Size=32B, Training Recipe Category=The SOTA Training Recipes, Rank=1%2026.07 | 78.88 | 93.33 | |
| PolarisModel Size=4B, Training Recipe Category=The SOTA Training Recipes, Type=RL2026.07 | 78.75 | 93.33 | |
| Projected (Ours)Model Size=4B, Training Recipe Category=The SOTA Training Recipes, Rank=10%2026.07 | 78.02 | 93.33 | |
| OLMo3-32B-think-dpoModel Size=32B, Training Recipe Category=The SOTA Training Recipes, Type=Base2026.07 | 74.48 | 93.33 | |
| Qwen3 4BModel Size=4B, Training Recipe Category=The SOTA Training Recipes, Type=Base2026.07 | 72.5 | 90 | |
| OLMo-3.1-7B-RL-Zero-MathModel Size=7B, Training Recipe Category=Emergent Reasoning Ability from the Base Model, Type=RL2026.07 | 50.21 | 76.67 | |
| Projected (Ours)Model Size=7B, Training Recipe Category=Emergent Reasoning Ability from the Base Model, Rank=30%2026.07 | 49.02 | 80 | |
| DeepscaleRModel Size=1.5B, Training Recipe Category=The SOTA Training Recipes, Type=RL2026.07 | 40.31 | 76.67 | |
| Projected (Ours)Model Size=1.5B, Training Recipe Category=The SOTA Training Recipes, Rank=1%2026.07 | 40.21 | 76.67 | |
| Deepseek-distill-qwen2.5Model Size=1.5B, Training Recipe Category=The SOTA Training Recipes, Type=Base2026.07 | 31.67 | 80 | |
| OLMo3-7B-BaseModel Size=7B, Training Recipe Category=Emergent Reasoning Ability from the Base Model, Type=Base2026.07 | 22.08 | 70 |