Mathematics on Base Aggregate Math (test)
69.7ScoreOlmo 3
Evaluation Results
| Method | Links | |
|---|---|---|
| Olmo 3Model Scale=32B, Training Stage=Stage 2 Soup, Cumulative training tokens=5.7T2025.12 | 69.7 | |
| Olmo 3Model Scale=32B, Training Stage=Stage 2 Ingredient 1, Cumulative training tokens=5.6T2025.12 | 66.8 | |
| Olmo 3Model Scale=32B, Training Stage=Stage 2 Ingredient 2, Cumulative training tokens=5.6T2025.12 | 65.4 | |
| Olmo 3Model Scale=32B, Training Stage=Stage 3, Cumulative training tokens=6.2T2025.12 | 61.4 | |
| Olmo 3Model Scale=7B, Training Stage=Stage 2, Cumulative training tokens=6T2025.12 | 59.8 | |
| Olmo 3Model Scale=7B, Training Stage=Stage 3, Cumulative training tokens=6.05T2025.12 | 54.4 | |
| OLMo 2Model Scale=32B, Training Stage=Stage 2 Soup, Cumulative training tokens=7.1T2025.12 | 53.9 | |
| OLMo 2Model Scale=32B, Training Stage=Stage 2 Ingredient 2, Cumulative training tokens=6.6T2025.12 | 51.9 | |
| OLMo 2Model Scale=32B, Training Stage=Stage 2 Ingredient 4, Cumulative training tokens=6.8T2025.12 | 51.9 | |
| OLMo 2Model Scale=32B, Training Stage=Stage 2 Ingredient 1, Cumulative training tokens=6.6T2025.12 | 51.6 | |
| OLMo 2Model Scale=32B, Training Stage=Stage 2 Ingredient 3, Cumulative training tokens=6.6T2025.12 | 51.5 | |
| MarinModel Scale=32B, Training Stage=Mantis, Cumulative training tokens=6.5T2025.12 | 49.3 | |
| Olmo 3Model Scale=32B, Training Stage=Stage 1, Cumulative training tokens=5.5T2025.12 | 48.4 | |
| K2Model Scale=70B, Training Stage=Stage 2, Cumulative training tokens=1.4T2025.12 | 43.3 | |
| OLMo 2Model Scale=7B, Training Stage=Stage 2 Soup, Cumulative training tokens=4.15T2025.12 | 41.7 | |
| OLMo 2Model Scale=7B, Training Stage=Stage 2 Ingredient 2, Cumulative training tokens=4.05T2025.12 | 41.4 | |
| OLMo 2Model Scale=7B, Training Stage=Stage 2 Ingredient 3, Cumulative training tokens=4.05T2025.12 | 40.8 | |
| ApertusModel Scale=70B, Training Stage=Phase 5, Cumulative training tokens=15T2025.12 | 40.6 | |
| MarinModel Scale=8B, Training Stage=Starling, Cumulative training tokens=12.4T2025.12 | 40.5 | |
| OLMo 2Model Scale=7B, Training Stage=Stage 2 Ingredient 1, Cumulative training tokens=4.05T2025.12 | 40.4 | |
| ApertusModel Scale=70B, Training Stage=Phase 4, Cumulative training tokens=13.5T2025.12 | 39.8 | |
| MarinModel Scale=8B, Training Stage=Deeper Starling, Cumulative training tokens=12.7T2025.12 | 39.4 | |
| ApertusModel Scale=70B, Training Stage=Phase 3, Cumulative training tokens=12T2025.12 | 34.2 | |
| K2Model Scale=70B, Training Stage=Stage 1, Cumulative training tokens=1.2T2025.12 | 34 | |
| OLMo 2Model Scale=32B, Training Stage=Stage 1, Cumulative training tokens=6.5T2025.12 | 33.2 | |
| ApertusModel Scale=8B, Training Stage=Phase 5, Cumulative training tokens=15T2025.12 | 29.3 | |
| ApertusModel Scale=8B, Training Stage=Phase 4, Cumulative training tokens=13.5T2025.12 | 26 | |
| MarinModel Scale=32B, Training Stage=Phase 3, Cumulative training tokens=5.4T2025.12 | 25.8 | |
| Olmo 3Model Scale=7B, Training Stage=Stage 1, Cumulative training tokens=5.9T2025.12 | 23.5 | |
| ApertusModel Scale=8B, Training Stage=Phase 3, Cumulative training tokens=12T2025.12 | 19.2 | |
| OLMo 2Model Scale=7B, Training Stage=Stage 1, Cumulative training tokens=4T2025.12 | 12.7 | |
| MarinModel Scale=8B, Training Stage=Phoenix, Cumulative training tokens=11.1T2025.12 | 11.2 |