Multi-task Language Understanding on MMLU (Accuracy, Exit Position)
73.1AccuracyMinistral3 8B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Ministral3 8BModel Architecture=Ministral3 8B (34 layers), Evaluation Protocol=Backbone2026.04 | 73.1 | — | |
| River-LLMModel Architecture=Ministral3 8B (34 layers), Early-exit threshold (tau)=0.8, Evaluation Protocol=Adaptive2026.04 | 72.7 | 2.01 | |
| River-LLMModel Architecture=Ministral3 8B (34 layers), Early-exit threshold (tau)=0.9, Evaluation Protocol=Adaptive2026.04 | 72.7 | 2.76 | |
| Phi4-miniModel Architecture=Phi4-mini (32 layers), Evaluation Protocol=Backbone2026.04 | 66.7 | — | |
| River-LLMModel Architecture=Phi4-mini (32 layers), Early-exit threshold (tau)=0.9, Evaluation Protocol=Adaptive2026.04 | 65.4 | 4.48 | |
| River-LLMModel Architecture=Phi4-mini (32 layers), Early-exit threshold (tau)=0.8, Evaluation Protocol=Adaptive2026.04 | 64.1 | 2.07 | |
| Pre-trainedModel=Gemma-3-4B2026.06 | 59.58 | — | |
| AlphaTokenModel=Gemma-3-4B2026.06 | 57.83 | — | |
| ssTOKENModel=Gemma-3-4B2026.06 | 56.62 | — | |
| STMModel=Gemma-3-4B2026.06 | 56.4 | — | |
| LoRAModel=Gemma-3-4B2026.06 | 55.92 | — | |
| XTFModel=Gemma-3-4B2026.06 | 55.8 | — | |
| Token CleaningModel=Gemma-3-4B2026.06 | 55.72 | — | |
| LESSModel=Gemma-3-4B2026.06 | 55.34 | — | |
| Standard FTModel=Gemma-3-4B2026.06 | 52.96 | — | |
| MoE with MPIModel Scale=11B, Setup=5-shot2026.06 | 50.93 | — | |
| MoEModel Scale=11B, Setup=5-shot2026.06 | 50 | — | |
| MoE with MPIModel Scale=3B, Setup=5-shot2026.06 | 48.83 | — | |
| MoEModel Scale=3B, Setup=5-shot2026.06 | 47.01 | — | |
| AuroraModel Size=1.1B, Evaluation Protocol=5-shot2026.06 | 38.8 | — | |
| NorMuonModel Size=1.1B, Evaluation Protocol=5-shot2026.06 | 30.1 | — | |
| MuonModel Size=1.1B, Evaluation Protocol=5-shot2026.06 | 29.7 | — | |
| U-NorMuonModel Size=1.1B, Evaluation Protocol=5-shot2026.06 | 29.1 | — | |
| NorMuonModel Size=340M, Evaluation Protocol=5-shot2026.06 | 25.7 | — | |
| U-NorMuonModel Size=340M, Evaluation Protocol=5-shot2026.06 | 25.4 | — | |
| MuonModel Size=340M, Evaluation Protocol=5-shot2026.06 | 24.7 | — | |
| AuroraModel Size=340M, Evaluation Protocol=5-shot2026.06 | 24 | — |