Knowledge on MMLU-Pro (test)
58.6AccuracyGAC + Token-φ
Evaluation Results
| Method | Links | |
|---|---|---|
| GAC + Token-φVariant=Token-φ2026.05 | 58.6 | |
| GAC w/o φVariant=w/o φ2026.05 | 57.8 | |
| HPTCategory=Recent hybrid SFT–RL2026.05 | 56.4 | |
| CHORDCategory=SFT–RL mixing baselines2026.05 | 56.2 | |
| LUFFYCategory=Recent hybrid SFT–RL2026.05 | 56 | |
| KL-ctrlCategory=Rule-based controllers2026.05 | 55.8 | |
| SRFTCategory=Recent hybrid SFT–RL2026.05 | 55.6 | |
| Nash-MTLCategory=Multi-objective solvers2026.05 | 55.4 | |
| GradNorm-ctrlCategory=Rule-based controllers2026.05 | 55.3 | |
| CAGradCategory=Multi-objective solvers2026.05 | 55.1 | |
| SFT-best + RLCategory=SFT–RL mixing baselines2026.05 | 51.3 | |
| GRPO (pure RL)Category=SFT–RL mixing baselines2026.05 | 45.8 | |
| DPOCategory=RL-free alignment2026.05 | 42.1 | |
| IPOCategory=RL-free alignment2026.05 | 41.5 | |
| SFT-best2026.05 | 38.4 | |
| Qwen2.5-7B-Inst.2026.05 | 24.7 | |
| Dense baselineModel Type=Dense, Compute=5.45e212025.06 | 14.12 | |
| MoE w/ optimal ARModel Type=MoE, Activation rate=20.07, Compute=2.86e21, Data reuse=-2025.06 | 13.59 | |
| MoE w/ optimal ARModel Type=MoE, Activation rate=20.07, Compute=2.86e21, Data reuse=strict2025.06 | 13.59 |