Visual Reasoning on REASONMAP Short questions
0.5998Weighted AccuracyGPT-5
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-52025.10 | 0.5998 | 0.0948 | |
| GPT-4o2025.10 | 0.4115 | 0.0684 | |
| Seed1.5-VL2025.10 | 0.342 | 0.0525 | |
| Qwen2.5-VL-7B-Instruct + REWARDMAPTraining Data=RPlustrain + Rtrain2025.10 | 0.3151 | 0.0621 | |
| Qwen2.5-VL-7B-Instruct + RL (baseline)Training Data=RPlustrain + Rtrain2025.10 | 0.2951 | 0.06 | |
| Qwen2.5-VL-7B-Instruct + SFT -> RLTraining Data=RPlustrain + Rtrain2025.10 | 0.2882 | 0.0588 | |
| Qwen2.5-VL-7B-Instruct + RL (REINFORCE++)Training Data=Rtrain2025.10 | 0.2717 | 0.0568 | |
| Qwen2.5-VL-72B-Instruct2025.10 | 0.2665 | 0.0509 | |
| Qwen2.5-VL-7B-Instruct + RL (GRPO)Training Data=Rtrain2025.10 | 0.2622 | 0.0552 | |
| Qwen2.5-VL-7B-Instruct + RL (ReMax)Training Data=Rtrain2025.10 | 0.2622 | 0.0557 | |
| Qwen2.5-VL-32B-Instruct2025.10 | 0.1649 | 0.0388 | |
| Qwen2.5-VL-7B-Instruct + SFTTraining Data=RPlustrain2025.10 | 0.1363 | 0.0409 | |
| Qwen2.5-VL-7B-Instruct2025.10 | 0.1328 | 0.0401 | |
| Kimi-VL-A3B-Instruct2025.10 | 0.1276 | 0.033 | |
| Qwen2.5-VL-3B-Instruct2025.10 | 0.0868 | 0.0275 | |
| Kimi-VL-A3B-Thinking2025.10 | 0.0547 | 0.0244 |