Multi-view Mathematical Reasoning on MVMATH
36.9AccuracyDAPO (Ours)
Evaluation Results
| Method | Links | |
|---|---|---|
| DAPO (Ours)Model Category=Multi-image/Video Enhancing MLLMs, Training Strategy=DAPO2026.01 | 36.9 | |
| DPS (Ours)Model Category=Multi-image/Video Enhancing MLLMs, Training Strategy=DPS2026.01 | 35.1 | |
| Two-stage RL (Ours)Model Category=Multi-image/Video Enhancing MLLMs, Training Strategy=Two-stage RL (DPS and annealing)2026.01 | 35 | |
| GPT-4oModel Category=Closed-Source MLLMs2026.01 | 32.1 | |
| Gemini-1.5-ProModel Category=Closed-Source MLLMs2026.01 | 29.1 | |
| TW-GRPOModel Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=7B2026.01 | 28.2 | |
| Qwen2.5-VLModel Category=Open-Source General MLLMs, Parameter Scale=7B2026.01 | 26.7 | |
| VideoRFTModel Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=7B2026.01 | 25.1 | |
| GPT-4VModel Category=Closed-Source MLLMs2026.01 | 24.5 | |
| LLaVA-OneVisionModel Category=Open-Source General MLLMs, Parameter Scale=7B2026.01 | 19.1 | |
| InternVL2.5Model Category=Open-Source General MLLMs, Parameter Scale=8B2026.01 | 18.8 | |
| mPLUG-Owl3Model Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=8B2026.01 | 18.7 | |
| Mantis-Idefics2Model Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=8B2026.01 | 15.5 | |
| LLaVA-NeXT-InterleaveModel Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=7B2026.01 | 14.7 |