Spatial Reasoning (Video) on VSI-Bench
79.2AccuracyHuman
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Human2025.10 | 79.2 | — | — | — | — | — | — | |
| Ground Truth Semantic MapDevice=Jetson Orion, Model=Qwen3VL-32B-4bit2026.02 | 68.3 | 55.2 | 80.5 | 65.1 | 53.6 | 87.4 | 45.2 | |
| MosaicThinkerDevice=Jetson Orion, Model=Qwen3VL-32B-4bit2026.02 | 67.1 | 53.2 | 79.3 | 64.4 | 51.8 | 77.2 | 45 | |
| InternVL3.5-38BParameters=38B2025.10 | 66.3 | — | — | — | — | — | — | |
| APCDevice=Jetson Orion, Model=Qwen3VL-32B-4bit2026.02 | 62.5 | 48.7 | 75.9 | 62.2 | 49.4 | 66.3 | 43.3 | |
| Scene Graph ReconstructionDevice=Jetson Orion, Model=Qwen3VL-32B-4bit2026.02 | 58.6 | 44.5 | 76.5 | 59.1 | 48 | 70.2 | 42.7 | |
| Video-CoTDevice=Jetson Orion, Model=Qwen3VL-32B-4bit2026.02 | 57.7 | 46.1 | 77.4 | 59.7 | 47.9 | 68.4 | 44.9 | |
| Direct InputDevice=Jetson Orion, Model=Qwen3VL-32B-4bit2026.02 | 56.4 | 47.7 | 76.3 | 58.7 | 47.2 | 67.4 | 44.1 | |
| ViLAVT-7BThinking Level=Chatting with Images2026.02 | 52 | — | — | — | — | — | — | |
| SpaceVista-7BParameters=7B, RL=true2025.10 | 48.6 | — | — | — | — | — | — | |
| Spatial-MLLM-4BThinking Level=Thinking about Images2026.02 | 48.4 | — | — | — | — | — | — | |
| SpatialMLLM-4BParameters=4B2025.10 | 48.4 | — | — | — | — | — | — | |
| SpaceR-7BParameters=7B2025.10 | 46.9 | — | — | — | — | — | — | |
| SpaceVista-7BParameters=7B2025.10 | 46.3 | — | — | — | — | — | — | |
| VG LLM-4BParameters=4B2025.10 | 46.1 | — | — | — | — | — | — | |
| SpaceR-7BThinking Level=Thinking about Images2026.02 | 45.6 | — | — | — | — | — | — | |
| VILASR-7BThinking Level=Thinking with Images2026.02 | 45.4 | — | — | — | — | — | — | |
| VILASR-7BParameters=7B2025.10 | 45.4 | — | — | — | — | — | — | |
| Gemini-2.5-pro2025.10 | 45 | — | — | — | — | — | — | |
| GPT-52025.10 | 44.2 | — | — | — | — | — | — | |
| InternVL3-8BThinking Level=Non-Thinking2026.02 | 42.1 | — | — | — | — | — | — | |
| Qwen2.5-VL-7BFine-tuned on=SpaceVista-1M2025.10 | 42 | — | — | — | — | — | — | |
| InternVL3.5-8BParameters=8B2025.10 | 38.2 | — | — | — | — | — | — | |
| LLaVA-NeXT-Video-7BParameters=7B2025.10 | 35.6 | — | — | — | — | — | — | |
| Qwen2.5-VL-7BThinking Level=Non-Thinking2026.02 | 34.7 | — | — | — | — | — | — | |
| Qwen2.5-VL-7BParameters=7B2025.10 | 32.7 | — | — | — | — | — | — | |
| LLaVA-OneVision-7BThinking Level=Non-Thinking2026.02 | 32.4 | — | — | — | — | — | — | |
| LLAVA-Onevision-7BParameters=7B2025.10 | 32.4 | — | — | — | — | — | — | |
| Qwen2.5-VL-72BParameters=72B2025.10 | 30.7 | — | — | — | — | — | — | |
| Qwen2.5-VL-7BThinking Level=Thinking about Images2026.02 | 26.2 | — | — | — | — | — | — |