Multimodal Reasoning on MathVista, WeMath, ChartQA, LogicVista, MMStar, VisPuzzles, and RealWorldQA
78.5MathVista AccuracyROMA
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| ROMAParameters=8B, Optimization=ROMA, Base Model=Qwen3-VL-8B Instruct2026.05 | 78.5 | 77.9 | 80.8 | 62.1 | 69.5 | 42.5 | 69.9 | 68.7 | |
| Qwen3-VL-8B Instruct + GRPOParameters=8B, Optimization=GRPO2026.05 | 78.4 | 77.6 | 81.5 | 60.8 | 70.1 | 43.5 | 70.6 | 68.9 | |
| Qwen3-VL-8B InstructParameters=8B2026.05 | 76.6 | 69.4 | 79.4 | 60.7 | 68.7 | 43.5 | 69.4 | 66.8 | |
| PAPO-7BParameters=7B2026.05 | 75.5 | 71 | 82 | 53.3 | 63.2 | 35.8 | 67.3 | 64 | |
| NoisyRollout-7BParameters=7B2026.05 | 72.7 | 69.3 | 79.8 | 50 | 63.2 | 37.7 | 67.1 | 62.8 | |
| VL-Rethinker-7BParameters=7B2026.05 | 72.7 | 67.5 | 79.9 | 46.9 | 61.9 | 34.8 | 68.5 | 61.7 | |
| Vision-R1-7BParameters=7B2026.05 | 72.4 | — | 81.6 | 48.7 | 62.7 | 36.1 | 66.1 | — | |
| OpenVLThinker-7BParameters=7B2026.05 | 67 | 60.6 | 78.7 | 48 | 60.1 | 32 | 58.6 | 57.9 |