Multi-task Language Understanding on MMLU Pro (Accuracy)
96.8AccuracyPass@100
Evaluation Results
| Method | Links | |
|---|---|---|
| Pass@100Generator Model=Llama 3.1 8B Instruct, Supervision Setting=Oracle2026.04 | 96.8 | |
| Pass@100Generator Model=Llama 3.3 70B Instruct, Supervision Setting=Oracle2026.04 | 92 | |
| LogisticGenerator Model=Llama 3.3 70B Instruct, Supervision Setting=Supervised2026.04 | 91.8 | |
| FUSEGenerator Model=Llama 3.3 70B Instruct, Supervision Setting=Unsupervised2026.04 | 91.4 | |
| Naive BayesGenerator Model=Llama 3.3 70B Instruct, Supervision Setting=Supervised2026.04 | 91.4 | |
| WeaverGenerator Model=Llama 3.3 70B Instruct, Supervision Setting=Supervised2026.04 | 90.2 | |
| Naive EnsembleGenerator Model=Llama 3.3 70B Instruct, Supervision Setting=Unsupervised2026.04 | 87 | |
| OBV (Oracle Best Verifier)Generator Model=Llama 3.3 70B Instruct, Supervision Setting=Supervised2026.04 | 85.6 | |
| Qwen3-30B-A3BModel Scale=30B, Active Parameters=3.3A, Teacher Model Status=False2026.05 | 81.11 | |
| NanoV3-30B*Model Scale=30B, Active Parameters=3.6A, Teacher Model Status=True2026.05 | 78.86 | |
| NanoV3 Elastic-30BModel Scale=30B, Active Parameters=3.6A, Teacher Model Status=False2026.05 | 78.63 | |
| NanoV3 Elastic-23BModel Scale=23B, Active Parameters=2.8A, Teacher Model Status=False2026.05 | 76.07 | |
| Majority VoteGenerator Model=Llama 3.3 70B Instruct, Supervision Setting=Unsupervised2026.04 | 74.4 | |
| Pass@1Generator Model=Llama 3.3 70B Instruct, Supervision Setting=Unsupervised2026.04 | 69.9 | |
| FUSEGenerator Model=Llama 3.1 8B Instruct, Supervision Setting=Unsupervised2026.04 | 69.8 | |
| NanoV3 Elastic-12BModel Scale=12B, Active Parameters=2.0A, Teacher Model Status=False2026.05 | 68.28 | |
| Naive EnsembleGenerator Model=Llama 3.1 8B Instruct, Supervision Setting=Unsupervised2026.04 | 67.2 | |
| WeaverGenerator Model=Llama 3.1 8B Instruct, Supervision Setting=Supervised2026.04 | 67.2 | |
| OBV (Oracle Best Verifier)Generator Model=Llama 3.1 8B Instruct, Supervision Setting=Supervised2026.04 | 66.5 | |
| Llama 3.1 InstructModel Scale=70B2025.04 | 66.3 | |
| Naive BayesGenerator Model=Llama 3.1 8B Instruct, Supervision Setting=Supervised2026.04 | 66 | |
| LogisticGenerator Model=Llama 3.1 8B Instruct, Supervision Setting=Supervised2026.04 | 65.2 | |
| Llama 3 InstructModel Scale=70B2025.04 | 63.2 | |
| Qwen3 ModelParameters=8B2025.09 | 63.2 | |
| ParamΔModel Scale=70B2025.04 | 62.1 | |
| SPELLBase Model=Qwen2.5-32B2025.09 | 60.22 | |
| SPELLBase Model=Qwen2.5-14B2025.09 | 58.86 | |
| Majority VoteGenerator Model=Llama 3.1 8B Instruct, Supervision Setting=Unsupervised2026.04 | 56.4 | |
| AAPAParameters=4B2025.09 | 56.35 | |
| Llama 3 BaseModel Scale=70B2025.04 | 54 | |
| Llama 3.1 BaseModel Scale=70B2025.04 | 51.3 | |
| SPELLBase Model=Qwen2.5-7B2025.09 | 49.78 | |
| Qwen2.5-32B2025.09 | 48.89 | |
| Llama 3.1 InstructModel Scale=8B2025.04 | 48.6 | |
| Qwen2.5-14B2025.09 | 46.67 | |
| Pass@1Generator Model=Llama 3.1 8B Instruct, Supervision Setting=Unsupervised2026.04 | 46.6 | |
| Llama 3 InstructModel Scale=8B2025.04 | 45.5 | |
| ParamΔModel Scale=8B2025.04 | 45.5 | |
| Qwen2.5-7B2025.09 | 40.24 | |
| Qwen3 ModelParameters=235B-25072025.09 | 39.45 | |
| Llama 3.1 BaseModel Scale=8B2025.04 | 36.4 | |
| Qwen3 ModelParameters=4B2025.09 | 35.56 | |
| Qwen3 ModelParameters=4B-25072025.09 | 33.1 | |
| Llama 3 BaseModel Scale=8B2025.04 | 33 | |
| AAPAParameters=0.6B2025.09 | 27.08 | |
| Qwen3 ModelParameters=0.6B2025.09 | 23.63 | |
| Qwen3 ModelParameters=1.7B2025.09 | 20.53 | |
| Qwen3 ModelParameters=32B2025.09 | 20.13 | |
| DenseBase Model=Llama-3.2-3B, Model architecture type=Dense, Few-shot=true2025.06 | 19.6 | |
| Reg. MoE# train tokens=1T, Active parameters=1B2026.05 | 19.3 | |
| MoBBase Model=Llama-3.2-3B, Model architecture type=MoB, Few-shot=true2025.06 | 19.1 | |
| MiCRoBase Model=Llama-3.2-3B, Model architecture type=MiCRo, Few-shot=true2025.06 | 19 | |
| OLMoE# train tokens=5T, Active parameters=1B2026.05 | 18.7 | |
| EMO# train tokens=1T, Active parameters=1B2026.05 | 18.5 | |
| Reg. MoE# train tokens=130B, Active parameters=1B2026.05 | 15.8 | |
| EMO# train tokens=130B, Active parameters=1B2026.05 | 15.5 | |
| Dense# train tokens=130B, Active parameters=1B2026.05 | 12.2 | |
| DenseBase Model=Llama-3.2-1B, Model architecture type=Dense, Few-shot=true2025.06 | 11.2 | |
| MoBBase Model=Llama-3.2-1B, Model architecture type=MoB, Few-shot=true2025.06 | 11 | |
| MiCRoBase Model=Llama-3.2-1B, Model architecture type=MiCRo, Few-shot=true2025.06 | 10.7 | |
| MiCRoBase Model=SmollM2-360M, Model architecture type=MiCRo, Few-shot=true2025.06 | 10.1 | |
| DenseBase Model=SmollM2-360M, Model architecture type=Dense, Few-shot=true2025.06 | 9.9 | |
| MoBBase Model=SmollM2-360M, Model architecture type=MoB, Few-shot=true2025.06 | 9.8 | |
| MiCRoBase Model=SmollM2-135M, Model architecture type=MiCRo, Few-shot=true2025.06 | 7.9 | |
| DenseBase Model=SmollM2-135M, Model architecture type=Dense, Few-shot=true2025.06 | 7.8 | |
| MoBBase Model=SmollM2-135M, Model architecture type=MoB, Few-shot=true2025.06 | 7.4 |