Multi-view spatial reasoning on MINDCUBE-1k
62.35Overall AccuracypySpatial
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| pySpatial2026.03 | 62.35 | 41.83 | 64.89 | 72.67 | |
| GPT-4oReference=OpenAI (2024)2026.03 | 42.29 | 35 | 43 | 46.4 | |
| Cognitive MapBackbone=Qwen2.5-VL-3B-Instruct2026.03 | 41.43 | 37 | 41.67 | 44.4 | |
| VADARReference=Marsili et al. (2025)2026.03 | 40.76 | 33.5 | 40.67 | 46.8 | |
| Chain-of-ThoughtBackbone=Qwen2.5-VL-3B-Instruct2026.03 | 40.48 | 32 | 36 | 58 | |
| Qwen2.5-VL-3B-InstructReference=Bai et al. (2025)2026.03 | 37.81 | 34 | 36 | 45.2 | |
| View InterpolationReference=Yin et al. (2025), Backbone=Qwen2.5-VL-3B-Instruct2026.03 | 37.81 | 35.5 | 36.5 | 42.8 | |
| ViperGPTReference=Surís et al. (2023)2026.03 | 36.95 | 20.5 | 41 | 40.4 | |
| VADAR w/ Recon.Module=3D reconstruction2026.03 | 35.62 | 31 | 36.83 | 36.4 |