Reasoning on MATH-500 (Accuracy, Spikes, Repairs)
94.5Accuracy (%)CoPD
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| CoPDTraining Strategy=Co-Evolving Policy Distillation2026.04 | 94.5 | — | — | — | — | — | |
| Image-ExpertTraining Strategy=Image-Specific Expert2026.04 | 93.9 | — | — | — | — | — | |
| Text-ExpertTraining Strategy=Text-Specific Expert2026.04 | 93.65 | — | — | — | — | — | |
| Mixed RLVRTraining Strategy=Joint optimization on combined budget2026.04 | 93.55 | — | — | — | — | — | |
| OPD V→TTraining Strategy=Static On-policy Policy Distillation (Image to Text)2026.04 | 93.45 | — | — | — | — | — | |
| OPD T→VTraining Strategy=Static On-policy Policy Distillation (Text to Image)2026.04 | 93.4 | — | — | — | — | — | |
| BaseTraining Strategy=Base Model2026.04 | 92.8 | — | — | — | — | — | |
| BaselineBase Model=Qwen3-8B2026.04 | 71.2 | 500 | 356 | — | — | — | |
| SPREGBase Model=Qwen3-8B2026.04 | 71 | 500 | 355 | -0.2 | 4,559 | 4,559 |