Multilingual General Knowledge on Global MMLU Lite (18 languages)
90.13AccuracyNemotron-3-Ultra 550B-A55B-Base
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Nemotron-3-Ultra 550B-A55B-Baseshots=5-shot2026.06 | 90.13 | — | |
| Mistral-Large-3 675B-Base-2512shots=5-shot2026.06 | 87.34 | — | |
| GLM-4.5 Baseshots=5-shot2026.06 | 85.81 | — | |
| Kimi-K2 Baseshots=5-shot2026.06 | 85.63 | — | |
| DeepSeek-V3.2 Exp-Baseshots=5-shot2026.06 | 85.59 | — | |
| Qwen2.5-7B-Instruct + Translate TestTraining Stage=Instruct, Evaluation Protocol=Translate Test2026.01 | 53.73 | 96.49 | |
| TinyAya GlobalScenario=Cultural Adaptation of a Baseline SFT Model2026.06 | 53.5 | — | |
| Qwen2.5-7B + RLVRTraining Stage=SFT + RLVR2026.01 | 53.15 | 99.78 | |
| SP3F-7BTraining Stage=Full Pipeline2026.01 | 50.76 | 99.45 | |
| Qwen2.5-7B-InstructTraining Stage=Instruct2026.01 | 48.2 | 96.21 | |
| SFT (Full MDolci)Scenario=Marker-Augmented Finetuning of a Base Model2026.06 | 44.9 | — | |
| Qwen2.5-7B + SFTTraining Stage=SFT2026.01 | 13.48 | 89.62 | |
| Qwen2.5-7BTraining Stage=Base2026.01 | 8.34 | 85.85 | |
| Marker-Augmented FinetuningScenario=Marker-Augmented Finetuning of a Base Model, Base Model=SFT (Full MDolci), Markers=Enabled2026.06 | 2.3 | — | |
| Cultural SFTScenario=Cultural Adaptation of a Baseline SFT Model, Base Model=TinyAya Global2026.06 | -1.6 | — |