Spatial Reasoning (Multi-Image) on SPAR-Bench
67.3AccuracyHuman
Evaluation Results
| Method | Links | |
|---|---|---|
| Human2025.10 | 67.3 | |
| SpatialClawBackbone=Gemma4-31B2026.06 | 63.3 | |
| SpaceTools ToolshedBackbone=Gemma4-31B2026.06 | 53.9 | |
| ViLAVT-7BThinking Level=Chatting with Images2026.02 | 52.6 | |
| pySpatialBackbone=Gemma4-31B2026.06 | 51.7 | |
| SpaceVista-7BParameters=7B, RL=true2025.10 | 41.6 | |
| AdaTooler-V-7BParameters=7B2025.12 | 40.3 | |
| SpaceVista-7BParameters=7B2025.10 | 38.1 | |
| VILASR-7BThinking Level=Thinking with Images2026.02 | 37.6 | |
| VILASRCategory=Open-Source o3-like2025.12 | 37.6 | |
| SpaceR-7BParameters=7B2025.10 | 37.6 | |
| VILASR-7BParameters=7B2025.10 | 37.6 | |
| GPT-52025.10 | 37.4 | |
| SpaceR-7BThinking Level=Thinking about Images2026.02 | 37.1 | |
| Qwen2.5-VL-7BThinking Level=Non-Thinking2026.02 | 36.9 | |
| Qwen2.5-VL-7BFine-tuned on=SpaceVista-1M2025.10 | 36.9 | |
| Gemini-2.5-pro2025.10 | 36.3 | |
| InternVL3-8BThinking Level=Non-Thinking2026.02 | 36 | |
| InternVL3.5-8BParameters=8B2025.10 | 36 | |
| Spatial-MLLM-4BThinking Level=Thinking about Images2026.02 | 35.1 | |
| GPT-4oCategory=Proprietary2025.12 | 33.6 | |
| Qwen2.5-VL-7BParameters=7B2025.10 | 33.1 | |
| Qwen2.5-VL-7B-InstructParameters=7B2025.12 | 33.07 | |
| LLaVA-OneVision-7BThinking Level=Non-Thinking2026.02 | 32.4 | |
| Qwen2.5-VL-72BParameters=72B2025.10 | 32.4 | |
| Qwen2.5-VL-7BThinking Level=Thinking about Images2026.02 | 31.6 | |
| SpatialMLLM-4BParameters=4B2025.10 | 31.5 | |
| LLaVA-NeXT-Video-7BParameters=7B2025.10 | 31.3 | |
| InternVL3.5-38BParameters=38B2025.10 | 31 | |
| LLAVA-Onevision-7BParameters=7B2025.10 | 30.6 | |
| LLaVA-1.5-7BParameters=7B2025.12 | 23.65 |