Image Reasoning on MathVision
57.88AccuracyCoPD
Evaluation Results
| Method | Links | |
|---|---|---|
| CoPDTraining Strategy=Co-Evolving Policy Distillation2026.04 | 57.88 | |
| CoPDTraining Strategy=CoPD2026.04 | 57.57 | |
| DAPO + ReMindBase Model=Qwen3-VL-8B-Instruct2026.06 | 57.32 | |
| MOPDTraining Strategy=MOPD2026.04 | 56.99 | |
| Mixed RLVRTraining Strategy=Joint optimization on combined budget2026.04 | 56.96 | |
| GRPO + ReMindBase Model=Qwen3-VL-8B-Instruct2026.06 | 56.81 | |
| OPD T→VTraining Strategy=Static On-policy Policy Distillation (Text to Image)2026.04 | 56.53 | |
| Mixed RLVRTraining Strategy=Mixed RLVR2026.04 | 56 | |
| Text-ExpertTraining Strategy=Text-Specific Expert2026.04 | 55.82 | |
| OPD V→TTraining Strategy=Static On-policy Policy Distillation (Image to Text)2026.04 | 55.82 | |
| Text-ExpertTraining Strategy=Text-Expert2026.04 | 55.82 | |
| Image-ExpertTraining Strategy=Image-Specific Expert2026.04 | 55.69 | |
| Image-ExpertTraining Strategy=Image-Expert2026.04 | 55.69 | |
| ExGRPOBase Model=Qwen3-VL-8B-Instruct2026.06 | 55.46 | |
| Video-ExpertTraining Strategy=Video-Expert2026.04 | 55.3 | |
| BaseTraining Strategy=Base Model2026.04 | 55.07 | |
| BaseTraining Strategy=Base2026.04 | 55.07 | |
| RLEPBase Model=Qwen3-VL-8B-Instruct2026.06 | 54.23 | |
| RePOBase Model=Qwen3-VL-8B-Instruct2026.06 | 53.77 | |
| DAPOBase Model=Qwen3-VL-8B-Instruct2026.06 | 52.78 | |
| GRPOBase Model=Qwen3-VL-8B-Instruct2026.06 | 48.82 | |
| Base ModelBase Model=Qwen3-VL-8B-Instruct2026.06 | 47.37 |