Mathematical Reasoning on OMEGA
53.4ScoreOlmo 3.1 Think 32B
Evaluation Results
| Method | Links | |
|---|---|---|
| Olmo 3.1 Think 32BTraining Stage=Final Think 3.1, Model Family=Olmo 3.1, Parameter Count=32B, Thinking Capability=true2025.12 | 53.4 | |
| Qwen 3 VL 32B ThinkModel Family=Qwen 3, Parameter Count=32B, Thinking Capability=true2025.12 | 50.8 | |
| Olmo 3 Think (Final 3.0)Training Stage=Final Think 3.0, Model Family=Olmo 3, Parameter Count=32B, Thinking Capability=true2025.12 | 50.6 | |
| Qwen 3 32BModel Family=Qwen 3, Parameter Count=32B, Thinking Capability=false2025.12 | 47.7 | |
| K2-V2 70B InstructModel Family=K2, Parameter Count=70B, Thinking Capability=false2025.12 | 46.1 | |
| Olmo 3 Think (DPO)Training Stage=DPO, Model Family=Olmo 3, Parameter Count=32B, Thinking Capability=true2025.12 | 45.2 | |
| Olmo 3 7B ThinkStage=Final Think2025.12 | 45 | |
| Qwen 3 8B2025.12 | 43.4 | |
| OR Nemotron 7B2025.12 | 43.2 | |
| Olmo 3 Think (SFT)Training Stage=SFT, Model Family=Olmo 3, Parameter Count=32B, Thinking Capability=true2025.12 | 43.1 | |
| Nemotron Nano 9B v22025.12 | 42.4 | |
| Olmo 3 7B ThinkStage=DPO2025.12 | 40.5 | |
| DS-R1 32BModel Family=DeepSeek-R1, Parameter Count=32B, Thinking Capability=true2025.12 | 38.9 | |
| OpenThinker3 7B2025.12 | 38.4 | |
| Qwen 3 VL 8B Think2025.12 | 38.1 | |
| Olmo 3 7B ThinkStage=SFT2025.12 | 37.8 | |
| Qwen 3 VL 8B Inststage=Instruct2025.12 | 32.3 | |
| Olmo 3 7B Instructstage=Final Instruct2025.12 | 28.9 | |
| DS-R1 Qwen 7B2025.12 | 28.5 | |
| Olmo 3 7B Instructstage=DPO2025.12 | 22.8 | |
| Qwen 3 8Bstage=Instruct2025.12 | 20.5 | |
| Gemma 3 12BParameters=12B2026.02 | 15.19 | |
| DictaLM 3.0 12B-InstParameters=12B, Variant=Instruct2026.02 | 15.19 | |
| Olmo 3 7B Instructstage=SFT2025.12 | 14.4 | |
| Qwen 2.5 7Bstage=Instruct2025.12 | 13.7 | |
| Granite 3.3 8B Inststage=Instruct2025.12 | 10.7 | |
| OLMo 2 7B Inststage=Instruct2025.12 | 5.2 | |
| Apertus 8B Inststage=Instruct2025.12 | 5 |