Image Reasoning on MMMU
67.5AccuracyOPD T→V
Evaluation Results
| Method | Links | |
|---|---|---|
| OPD T→VTraining Strategy=Static On-policy Policy Distillation (Text to Image)2026.04 | 67.5 | |
| CoPDTraining Strategy=Co-Evolving Policy Distillation2026.04 | 66.94 | |
| Mixed RLVRTraining Strategy=Joint optimization on combined budget2026.04 | 66.39 | |
| OPD V→TTraining Strategy=Static On-policy Policy Distillation (Image to Text)2026.04 | 65.94 | |
| Text-ExpertTraining Strategy=Text-Specific Expert2026.04 | 65.72 | |
| Image-ExpertTraining Strategy=Image-Specific Expert2026.04 | 65.28 | |
| BaseTraining Strategy=Base Model2026.04 | 65.06 | |
| Scaled base VLM 7B + Joint FFTTrain Data=3.6M + 1.6M2026.06 | 47.4 | |
| Scaled base VLM 7BTrain Data=3.6M2026.06 | 45.1 | |
| Scaled base VLM 7B + MERIT (Proposed, 2D)Train Data=3.6M + 1.6M2026.06 | 44.8 | |
| Base VLM 7B + Joint FFTTrain Data=0.7M + 1.6M2026.06 | 36.4 | |
| Base VLM 7B + MERIT (Proposed, 2D)Train Data=0.7M + 1.6M2026.06 | 36.1 | |
| LLaVA-1.5-7BTrain Data=0.7M2026.06 | 35.3 | |
| LLaVA-1.5-13B + Joint FFTTrain Data=0.7M + 0.2M2026.06 | 34.4 | |
| LLaVA-7BTrain Data=0.6M2026.06 | 34.1 | |
| LLaVA-1.5-13BTrain Data=0.7M2026.06 | 33.6 | |
| Base VLM 7BTrain Data=0.7M2026.06 | 32.8 |