General language understanding and reasoning on MMLU-Redux
83.7AccuracyQwen 3 14B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen 3 14Bshots=5-shot2026.01 | 83.7 | |
| Ministral 3 14Bshots=5-shot2026.01 | 82 | |
| Qwen 3 8Bshots=5-shot2026.01 | 79.4 | |
| Ministral 3 8Bshots=5-shot2026.01 | 79.3 | |
| Gemma 3 12Bshots=5-shot2026.01 | 76.6 | |
| Qwen 3 4Bshots=5-shot2026.01 | 75.9 | |
| Ministral 3 3Bshots=5-shot2026.01 | 73.5 | |
| HySparse# Shots=5-shot, Model Architecture=80B MoE (Hybrid 1:11), Attention Variant=HySparse2026.02 | 66.2 | |
| Full-Attn# Shots=5-shot, Model Architecture=80B MoE (Hybrid 1:11), Attention Variant=Full-Attn2026.02 | 65.6 | |
| Gemma 3 4Bshots=5-shot2026.01 | 62.6 | |
| HySparse# Shots=5-shot, Model Architecture=7B Dense (Hybrid 1:3), Attention Variant=HySparse2026.02 | 61.6 | |
| Hybrid SWA# Shots=5-shot, Model Architecture=7B Dense (Hybrid 1:3), Attention Variant=Hybrid SWA2026.02 | 60.8 | |
| Full-Attn# Shots=5-shot, Model Architecture=7B Dense (Hybrid 1:3), Attention Variant=Full-Attn2026.02 | 59.6 | |
| Hybrid SWA# Shots=5-shot, Model Architecture=80B MoE (Hybrid 1:11), Attention Variant=Hybrid SWA2026.02 | 57.4 |