Reasoning on AIME 2025 (Accuracy)
49.58AIME 2025 AccuracyCoPD
Evaluation Results
| Method | Links | |
|---|---|---|
| CoPDTraining Strategy=Co-Evolving Policy Distillation2026.04 | 49.58 | |
| CoPDTraining Strategy=CoPD2026.04 | 49.38 | |
| Text-ExpertTraining Strategy=Text-Specific Expert2026.04 | 48.33 | |
| Text-ExpertTraining Strategy=Text-Expert2026.04 | 48.33 | |
| Image-ExpertTraining Strategy=Image-Specific Expert2026.04 | 47.5 | |
| Image-ExpertTraining Strategy=Image-Expert2026.04 | 47.5 | |
| Video-ExpertTraining Strategy=Video-Expert2026.04 | 47.29 | |
| BaseTraining Strategy=Base Model2026.04 | 46.88 | |
| BaseTraining Strategy=Base2026.04 | 46.88 | |
| MOPDTraining Strategy=MOPD2026.04 | 46.04 | |
| Mixed RLVRTraining Strategy=Mixed RLVR2026.04 | 45.62 | |
| OPD T→VTraining Strategy=Static On-policy Policy Distillation (Text to Image)2026.04 | 45.42 | |
| Mixed RLVRTraining Strategy=Joint optimization on combined budget2026.04 | 44.58 | |
| OPD V→TTraining Strategy=Static On-policy Policy Distillation (Image to Text)2026.04 | 43.54 | |
| VanillaModel=Llama2025.09 | 43.3 | |
| OjaKVModel=Llama2025.09 | 13 | |
| Eigen-NModel=Llama2025.09 | 0 | |
| StaticPCAModel=Llama2025.09 | 0 | |
| PaluModel=Llama2025.09 | 0 |