Spatial Reasoning on SpatialEval
70.81AccuracyREWARDMAP
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| REWARDMAPModel=Qwen2.5-VL-7B-Instruct, Training Strategy=REWARDMAP2025.10 | 70.81 | — | — | — | |
| Qwen2.5-VL-7B-Instruct + SFT -> RLModel=Qwen2.5-VL-7B-Instruct, Training Strategy=SFT -> RL Baseline2025.10 | 69.06 | — | — | — | |
| Qwen2.5-VL-32B-InstructModel=Qwen2.5-VL-32B-Instruct, Training Strategy=Reference2025.10 | 57.99 | — | — | — | |
| Qwen2.5-VL-7B-InstructModel=Qwen2.5-VL-7B-Instruct, Training Strategy=Base Model2025.10 | 57.3 | — | — | — | |
| Qwen2.5-VL-3B-InstructModel=Qwen2.5-VL-3B-Instruct, Training Strategy=Reference2025.10 | 55.04 | — | — | — | |
| Kimi-VL-A3B-InstructModel=Kimi-VL-A3B-Instruct, Training Strategy=Reference2025.10 | 52.64 | — | — | — | |
| Intern8BInter Head=after2026.03 | 48.45 | — | — | — | |
| Intern8BInter Head=before2026.03 | 47.92 | — | — | — | |
| Qwen7BInter Head=after2026.03 | 45.81 | — | — | — | |
| Qwen7BInter Head=before2026.03 | 45.33 | — | — | — | |
| Llama11BInter Head=after2026.03 | 39.9 | — | — | — | |
| Llama11BInter Head=before2026.03 | 39.61 | — | — | — | |
| LLaVA-1.5-7BZero-shot=true2025.02 | — | 28.4 | 28.8 | 41.6 | |
| LLaVA-1.6-7BZero-shot=true2025.02 | — | 28 | 34.8 | 32.2 | |
| Magma-8B (Act)Zero-shot=true, SoM/ToM pre-training=without2025.02 | — | 36.9 | 44.8 | 37.5 | |
| Magma-8B (Full)Zero-shot=true, SoM/ToM pre-training=without2025.02 | — | 27.5 | 33.5 | 47.3 | |
| Magma-8B (Full)Zero-shot=true, SoM/ToM pre-training=with2025.02 | — | 43.4 | 36.5 | 64.5 | |
| Qwen-VL-9.6BZero-shot=true2025.02 | — | 28.7 | 31.8 | 25.7 |