General Knowledge on MMLU (MMLU, UtilityNorm, Score, Rank)
89.08MMLU AccuracyNemotron-3-Ultra 550B-A55B-Base
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Nemotron-3-Ultra 550B-A55B-Baseshots=5-shot2026.06 | 89.08 | — | — | — | — | |
| DeepSeek-V3.2 Exp-Baseshots=5-shot2026.06 | 87.82 | — | — | — | — | |
| Kimi-K2 Baseshots=5-shot2026.06 | 87.6 | — | — | — | — | |
| Mistral-Large-3 675B-Base-2512shots=5-shot2026.06 | 87.35 | — | — | — | — | |
| GLM-4.5 Baseshots=5-shot2026.06 | 86.5 | — | — | — | — | |
| LLaDA + DreamEnsemble Type=Intermediate-generation Ensemble, Scoring Function=Token change count2026.06 | 67.55 | — | — | — | — | |
| LLaDA + DreamEnsemble Type=Intermediate-generation Ensemble, Scoring Function=Entropy2026.06 | 67.53 | — | — | — | — | |
| DreamEnsemble Type=Single Model2026.06 | 67.46 | — | — | — | — | |
| LLaDA + DreamEnsemble Type=Intermediate-generation Ensemble, Scoring Function=Probability margin2026.06 | 67.34 | — | — | — | — | |
| LLaDA + DreamEnsemble Type=Post-generation Ensemble2026.06 | 67.26 | — | — | — | — | |
| LLaDA + DreamEnsemble Type=Intermediate-generation Ensemble, Scoring Function=Top-1 probability2026.06 | 67.25 | — | — | — | — | |
| LLaDA-8BParadigm=Masked Diffusion, Training Tokens=2.3T, Training Data=Not Released, Evaluation protocol=Reported by prior work, Shots=52026.06 | 65.9 | — | — | — | — | |
| Llama 3-8BParadigm=AR, Training Tokens=15T, Training Data=Not Released, Evaluation protocol=Reported by prior work, Shots=52026.06 | 65.4 | — | — | — | — | |
| SDBackbone=Llama-3.1-8B/Llama-3.2-1B2025.09 | 65 | — | — | — | 4.36 | |
| AutoJudge-RBackbone=Llama-3.1-8B/Llama-3.2-1B, Thresholding strategy=best recall2025.09 | 64.6 | — | — | — | 4.47 | |
| SelfJudge-RBackbone=Llama-3.1-8B/Llama-3.2-1B, Thresholding strategy=best recall2025.09 | 64.4 | — | — | — | 5.14 | |
| OriginalBase Model=Ministral-8B-Instruct-24102025.09 | 64 | — | — | — | — | |
| AutoJudge-FBackbone=Llama-3.1-8B/Llama-3.2-1B, Thresholding strategy=best F1 score2025.09 | 63.1 | — | — | — | 5.15 | |
| SelfJudge-FBackbone=Llama-3.1-8B/Llama-3.2-1B, Thresholding strategy=best F1 score2025.09 | 62.7 | — | — | — | 6.38 | |
| LLaDAEnsemble Type=Single Model2026.06 | 61.23 | — | — | — | — | |
| DOORBase Model=Ministral-8B-Instruct-24102025.09 | 55.5 | 98.5 | 28.9 | 5 | — | |
| Task VectorBase Model=Ministral-8B-Instruct-24102025.09 | 53.3 | 100 | 0 | 8 | — | |
| Sumi-7BParadigm=Uniform Diffusion, Training Tokens=1.5T, Training Data=Fully Released, Evaluation protocol=Evaluated under our protocol, Shots=52026.06 | 51.1 | — | — | — | — | |
| Top-kBackbone=Llama-3.1-8B/Llama-3.2-1B, k=22025.09 | 49.5 | — | — | — | 6.69 | |
| GAGDRBase Model=Ministral-8B-Instruct-24102025.09 | 48 | 92.1 | 52.8 | 4 | — | |
| NPOGDRBase Model=Ministral-8B-Instruct-24102025.09 | 46.3 | 80.6 | 55.6 | 3 | — | |
| Llama 2-7BParadigm=AR, Training Tokens=2T, Training Data=Not Released, Evaluation protocol=Evaluated under our protocol, Shots=52026.06 | 46 | — | — | — | — | |
| ConSA (head-wise, single-layer)Target Sparsity (ρ)=0.50, Granularity=head-wise, Constraint Scope=single-layer, Model=1.7B2026.06 | 45.76 | — | — | — | — | |
| ConSA (head-wise, all-layers)Target Sparsity (ρ)=0.50, Granularity=head-wise, Constraint Scope=all-layers, Model=1.7B2026.06 | 45.55 | — | — | — | — | |
| Dense FATarget Sparsity (ρ)=0, Model=1.7B2026.06 | 45.51 | — | — | — | — | |
| ConSA (layer-wise)Target Sparsity (ρ)=0.50, Granularity=layer-wise, Model=1.7B2026.06 | 45.45 | — | — | — | — | |
| Rule (head-wise)Target Sparsity (ρ)=0.50, Granularity=head-wise, Model=1.7B2026.06 | 45.15 | — | — | — | — | |
| Rule (layer-wise)Target Sparsity (ρ)=0.50, Granularity=layer-wise, Model=1.7B2026.06 | 44.03 | — | — | — | — | |
| PRISMBase Model=Ministral-8B-Instruct-24102025.09 | 38.5 | 73.2 | 76.1 | 1 | — | |
| OLMo-7BParadigm=AR, Training Tokens=2.5T, Training Data=Fully Released, Evaluation protocol=Evaluated under our protocol, Shots=52026.06 | 28 | — | — | — | — | |
| Falcon-7BParadigm=AR, Training Tokens=1.5T, Training Data=Partially Released, Evaluation protocol=Evaluated under our protocol, Shots=52026.06 | 27.2 | — | — | — | — | |
| SAM+NPOBase Model=Ministral-8B-Instruct-24102025.09 | 23.8 | 53 | 72.1 | 2 | — | |
| NPOBase Model=Ministral-8B-Instruct-24102025.09 | 23 | 0.3 | 5.5 | 6 | — | |
| GABase Model=Ministral-8B-Instruct-24102025.09 | 22.9 | 0 | 0 | 7 | — | |
| Fullsetting=Train-from-scratch2026.04 | — | — | 44.99 | — | — | |
| Full AttentionSetting=CPT setting2026.04 | — | — | 71.83 | — | — | |
| Hybrid-GDNsetting=Train-from-scratch2026.04 | — | — | 46.23 | — | — | |
| Hybrid-KSASetting=CPT setting2026.04 | — | — | 70.5 | — | — | |
| Hybrid-KSAsetting=Train-from-scratch2026.04 | — | — | 46.83 | — | — | |
| Hybrid-LinearSetting=CPT setting2026.04 | — | — | 64.33 | — | — | |
| Hybrid-SCASetting=CPT setting2026.04 | — | — | 69.83 | — | — | |
| Hybrid-SCAsetting=Train-from-scratch2026.04 | — | — | 46.77 | — | — | |
| Hybrid-SWASetting=CPT setting2026.04 | — | — | 70.57 | — | — | |
| Hybrid-SWAsetting=Train-from-scratch2026.04 | — | — | 46.84 | — | — | |
| KSASetting=CPT setting2026.04 | — | — | 70.73 | — | — | |
| KSAsetting=Train-from-scratch2026.04 | — | — | 46.83 | — | — | |
| Mellum 2Parameters=2.5B/12B2026.05 | — | — | 70.9 | — | — | |
| OLMo-3-7BParameters=7B2026.05 | — | — | 62.1 | — | — | |
| Qwen2.5-7BParameters=7B2026.05 | — | — | 71.8 | — | — | |
| Qwen3-4BParameters=4B2026.05 | — | — | 71.1 | — | — | |
| Qwen3.5-4BParameters=4B2026.05 | — | — | 74.2 | — | — |