Visual Reasoning on REASONMAP Long questions
62.5Weighted AccuracyGPT-5
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-52025.10 | 62.5 | 19.75 | |
| GPT-4o2025.10 | 42.8 | 13.57 | |
| Seed1.5-VL2025.10 | 38.02 | 11.96 | |
| Qwen2.5-VL-7B-Instruct + REWARDMAPTraining Data=RPlustrain + Rtrain2025.10 | 31.77 | 11.22 | |
| Qwen2.5-VL-7B-Instruct + SFT -> RLTraining Data=RPlustrain + Rtrain2025.10 | 30.38 | 10.62 | |
| Qwen2.5-VL-7B-Instruct + RL (baseline)Training Data=RPlustrain + Rtrain2025.10 | 29.51 | 10.41 | |
| Qwen2.5-VL-7B-Instruct + RL (REINFORCE++)Training Data=Rtrain2025.10 | 27.6 | 10.12 | |
| Qwen2.5-VL-7B-Instruct + RL (ReMax)Training Data=Rtrain2025.10 | 27.26 | 9.99 | |
| Qwen2.5-VL-7B-Instruct + RL (GRPO)Training Data=Rtrain2025.10 | 26.04 | 9.52 | |
| Qwen2.5-VL-72B-Instruct2025.10 | 24.22 | 8.8 | |
| Qwen2.5-VL-32B-Instruct2025.10 | 15.71 | 6.84 | |
| Kimi-VL-A3B-Instruct2025.10 | 12.33 | 5.37 | |
| Qwen2.5-VL-7B-Instruct + SFTTraining Data=RPlustrain2025.10 | 9.11 | 6.25 | |
| Qwen2.5-VL-3B-Instruct2025.10 | 7.99 | 3.7 | |
| Qwen2.5-VL-7B-Instruct2025.10 | 7.12 | 5.74 | |
| Kimi-VL-A3B-Thinking2025.10 | 5.47 | 3.17 |