Multilingual Language Modeling on Multilingual MMMLU MGSM
55.67MMMLU ScoreGated Attention
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Gated AttentionArchitecture=24B-A3B, Precision=BF16, Training Tokens=500B2026.01 | 55.67 | 59.27 | |
| GatedNormArchitecture=24B-A3B, Precision=BF16, Training Tokens=500B2026.01 | 55.47 | 59.72 | |
| PreAffineArchitecture=24B-A3B, Precision=BF16, Training Tokens=500B2026.01 | 55.46 | 58.46 | |
| GatedNormArchitecture=24B-A3B, Precision=W4A4, Training Tokens=500B2026.01 | 53.7 | 56.78 | |
| GatedNormArchitecture=24B-A3B, Precision=W4A4+SQ, Training Tokens=500B2026.01 | 53.39 | 56 | |
| Gated AttentionArchitecture=24B-A3B, Precision=W4A4, Training Tokens=500B2026.01 | 53.21 | 51.85 | |
| Gated AttentionArchitecture=24B-A3B, Precision=W4A4+SQ, Training Tokens=500B2026.01 | 53.2 | 52.55 | |
| PreAffineArchitecture=24B-A3B, Precision=W4A4+SQ, Training Tokens=500B2026.01 | 51.98 | 50.72 | |
| PreAffineArchitecture=24B-A3B, Precision=W4A4, Training Tokens=500B2026.01 | 51.35 | 49.58 | |
| PreAffineArchitecture=MoE-7B-A2B, Precision=BF16, Training Tokens=1.2T2026.01 | 46.42 | 39.9 | |
| Gated AttentionArchitecture=MoE-7B-A2B, Precision=BF16, Training Tokens=1.2T2026.01 | 46.16 | 37.4 | |
| GatedNormArchitecture=MoE-7B-A2B, Precision=BF16, Training Tokens=1.2T2026.01 | 45.76 | 38.27 |