Multitask Language Understanding on Global MMLU-Lite
64.5AccuracyBYOL-nya
Evaluation Results
| Method | Links | |
|---|---|---|
| BYOL-nyaParameters=12B, Pre-training Type=CPT2026.01 | 64.5 | |
| SDARScale=8B, Inference Mode=Zero-shot, Decoding=Static (one token per step)2026.06 | 60.8 | |
| Gemma-3Parameters=12B, Pre-training Type=PT2026.01 | 60.75 | |
| OPDLMScale=8B, Inference Mode=Zero-shot, Decoding=Static (one token per step)2026.06 | 56 | |
| BYOL-nyaParameters=4B, Pre-training Type=CPT2026.01 | 55.25 | |
| OPDLMScale=4B, Inference Mode=Zero-shot, Decoding=Static (one token per step)2026.06 | 51.6 | |
| Fast-dLLM-v2Scale=7B, Inference Mode=Zero-shot, Decoding=Static (one token per step)2026.06 | 51.5 | |
| Gemma-3Parameters=4B, Pre-training Type=PT2026.01 | 50.75 | |
| SDARScale=4B, Inference Mode=Zero-shot, Decoding=Static (one token per step)2026.06 | 50.7 | |
| Qwen-3Parameters=14B, Pre-training Type=Base2026.01 | 45.75 | |
| Llama-3.1Parameters=8B, Pre-training Type=Base2026.01 | 43.25 | |
| ApertusParameters=8B, Pre-training Type=25092026.01 | 41.75 | |
| Qwen-3Parameters=8B, Pre-training Type=Base2026.01 | 37.5 | |
| Qwen-3Parameters=1.7B, Pre-training Type=Base2026.01 | 32.5 | |
| Llama-3.2Parameters=1B, Pre-training Type=Base2026.01 | 28.25 | |
| Gemma-3Parameters=1B, Pre-training Type=PT2026.01 | 26.75 | |
| BYOL-nyaParameters=1B, Pre-training Type=CPT2026.01 | 23 |