Multimodal Reasoning on Single-image benchmarks suite Overall
45.6Overall AccuracyOurs [with DPS and annealing]
Evaluation Results
| Method | Links | |
|---|---|---|
| Ours [with DPS and annealing]Model Category=Reasoning MLLMs, Training Strategy=Two-stage RL2026.01 | 45.6 | |
| OursModel Category=Reasoning MLLMs, Training Strategy=DAPO2026.01 | 45.5 | |
| Ours [with DPS]Model Category=Reasoning MLLMs, Training Strategy=DPS2026.01 | 45.4 | |
| VLAA-Thinker 7BModel Category=Reasoning MLLMs2026.01 | 44 | |
| VisonR1 7BModel Category=Reasoning MLLMs2026.01 | 43.4 | |
| MixedR1 7BModel Category=Reasoning MLLMs2026.01 | 43 | |
| R1-Onevision 7BModel Category=Reasoning MLLMs2026.01 | 41.4 | |
| Qwen2.5VL 7BModel Category=Open-Source General MLLMs2026.01 | 40.8 |