General Knowledge QA on MMLU Redux and MMLU Pro (test)
69.7MMLU-R ScoreGatedNorm
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GatedNormArchitecture=24B-A3B, Precision=BF16, Training Tokens=500B2026.01 | 69.7 | 47.13 | |
| PreAffineArchitecture=24B-A3B, Precision=BF16, Training Tokens=500B2026.01 | 67.96 | 48.48 | |
| Gated AttentionArchitecture=24B-A3B, Precision=BF16, Training Tokens=500B2026.01 | 67.49 | 46.02 | |
| GatedNormArchitecture=24B-A3B, Precision=W4A4+SQ, Training Tokens=500B2026.01 | 67.35 | 47.15 | |
| PreAffineArchitecture=24B-A3B, Precision=W4A4+SQ, Training Tokens=500B2026.01 | 67.17 | 43.44 | |
| GatedNormArchitecture=24B-A3B, Precision=W4A4, Training Tokens=500B2026.01 | 66.92 | 46.36 | |
| Gated AttentionArchitecture=24B-A3B, Precision=W4A4+SQ, Training Tokens=500B2026.01 | 66.63 | 44.86 | |
| PreAffineArchitecture=24B-A3B, Precision=W4A4, Training Tokens=500B2026.01 | 66.05 | 43.61 | |
| Gated AttentionArchitecture=24B-A3B, Precision=W4A4, Training Tokens=500B2026.01 | 65.77 | 45.71 | |
| GatedNormArchitecture=MoE-7B-A2B, Precision=BF16, Training Tokens=1.2T2026.01 | 61.71 | 33.46 | |
| PreAffineArchitecture=MoE-7B-A2B, Precision=BF16, Training Tokens=1.2T2026.01 | 61.59 | 32 | |
| Gated AttentionArchitecture=MoE-7B-A2B, Precision=BF16, Training Tokens=1.2T2026.01 | 59.59 | 31.91 |