Single-image spatial reasoning on CV-Bench
80.72D AccuracyGemini-3-Pro
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Gemini-3-ProCategory=Proprietary2026.02 | 80.7 | 91.2 | 85.9 | |
| GPT-5.2Category=Proprietary2026.02 | 77.9 | 90.4 | 84.2 | |
| Qwen2.5-VL-72BCategory=Open-Weight2026.02 | 77.1 | 84.2 | 80.7 | |
| Qwen2.5-VL-32BCategory=Open-Weight2026.02 | 76.8 | 84 | 80.4 | |
| GPT-4.1Category=Proprietary2026.02 | 75 | 86.9 | 80.9 | |
| Qwen2.5-VL-7BCategory=Qwen2.5-VL-7B Based2026.02 | 74.8 | 83.3 | 79.1 | |
| InternVL-2.5-8BCategory=Open-Weight2026.02 | 74 | 79.5 | 76.8 | |
| InternVL-2.5-4BCategory=Open-Weight2026.02 | 73.2 | 74.6 | 73.9 | |
| SpatialLadder-3BCategory=Qwen2.5-VL-3B Based2026.02 | 72.2 | 74.6 | 73.4 | |
| SpaceR-7BCategory=Qwen2.5-VL-7B Based2026.02 | 71.9 | 81.1 | 76.4 | |
| Video-R1Category=Qwen2.5-VL-7B Based2026.02 | 71.7 | 74.9 | 73.3 | |
| LLaVA-OneVision-4BCategory=Open-Weight2026.02 | 70.8 | 79.2 | 75 | |
| HATCHCategory=Qwen2.5-VL-3B Based, Prompting=Direct answer without generating viewpoint-transition actions2026.02 | 70.5 | 78.4 | 74.5 | |
| Qwen2.5-VL-3BCategory=Qwen2.5-VL-3B Based2026.02 | 69.3 | 72.2 | 70.7 | |
| Spatial-MLLM-4BCategory=Qwen2.5-VL-3B Based2026.02 | 65.7 | 69.5 | 67.6 |