Visual Reasoning on REASONMAP-PLUS
88.95Weighted AccuracyGPT-5
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GPT-52025.10 | 88.95 | 86.4 | 91.6 | |
| Qwen2.5-VL-7B-Instruct + REWARDMAPTraining Data=RPlustrain + Rtrain2025.10 | 74.25 | 72.18 | 76.42 | |
| Seed1.5-VL2025.10 | 73.58 | 65.26 | 82.23 | |
| Qwen2.5-VL-7B-Instruct + RL (baseline)Training Data=RPlustrain + Rtrain2025.10 | 67.61 | 68.37 | 66.82 | |
| GPT-4o2025.10 | 64.42 | 59.28 | 69.77 | |
| Qwen2.5-VL-7B-Instruct + SFT -> RLTraining Data=RPlustrain + Rtrain2025.10 | 60.53 | 55.38 | 65.9 | |
| Qwen2.5-VL-32B-Instruct2025.10 | 58.32 | 46.96 | 70.14 | |
| Qwen2.5-VL-7B-Instruct + SFTTraining Data=RPlustrain2025.10 | 57.93 | 50.73 | 65.44 | |
| Qwen2.5-VL-72B-Instruct2025.10 | 53.21 | 43.46 | 63.36 | |
| Qwen2.5-VL-7B-Instruct + RL (ReMax)Training Data=Rtrain2025.10 | 45.39 | 38.37 | 52.7 | |
| Qwen2.5-VL-7B-Instruct + RL (GRPO)Training Data=Rtrain2025.10 | 44.64 | 37.57 | 52.01 | |
| Qwen2.5-VL-7B-Instruct + RL (REINFORCE++)Training Data=Rtrain2025.10 | 44.64 | 36.82 | 52.79 | |
| Qwen2.5-VL-7B-Instruct2025.10 | 44.21 | 37.39 | 51.32 | |
| Qwen2.5-VL-3B-Instruct2025.10 | 37.61 | 22.68 | 53.16 | |
| Kimi-VL-A3B-Thinking2025.10 | 33.95 | 18.17 | 50.39 | |
| Kimi-VL-A3B-Instruct2025.10 | 32.55 | 14.75 | 51.08 |