3D Spatial Reasoning on 3DSRBench
95.7AccuracyHuman Estimate
Evaluation Results
| Method | Links | |
|---|---|---|
| Human Estimate2025.04 | 95.7 | |
| GPT-5-0807Model Type=Proprietary2025.11 | 68.2 | |
| Emb. TextProtocol=CoT2026.01 | 67 | |
| Rotation TokensProtocol=Direct2026.01 | 67 | |
| Qwen3.5Parameters=35BA3B, Internal Reasoning (Think mode)=false2026.05 | 66.6 | |
| 3DThinker2025.10 | 65.6 | |
| SenseNova-U1Parameters=8B, Internal Reasoning (Think mode)=true2026.05 | 64.88 | |
| Emb. Tokens (ViTPose)Protocol=CoT2026.01 | 64 | |
| Rotation TextProtocol=CoT2026.01 | 64 | |
| SenseNova-U1Parameters=30BA3B, Internal Reasoning (Think mode)=true2026.05 | 62.96 | |
| w/ Entropy-RegularizationEvaluation Protocol=RL Fine-tuning2026.05 | 62.9 | |
| SpatialThinker-30B (Ours)Model Scale=30B, Training Dataset=STVQA-7K2025.11 | 62.1 | |
| Claude-4-Sonnet-0514Model Type=Proprietary2025.11 | 61.9 | |
| Qwen3-VL-30BModel Scale=30B, Model Type=Open-Source General2025.11 | 60.4 | |
| SpatialReasoner2025.10 | 60.3 | |
| SpatialReasonerType=Specialist Reference2026.05 | 60.3 | |
| Rotation TextProtocol=Direct2026.01 | 60 | |
| w/ Tool-Encourage RewardEvaluation Protocol=RL Fine-tuning2026.05 | 59.9 | |
| vanilla RFT (Base)Evaluation Protocol=RL Fine-tuning2026.05 | 59.2 | |
| Emb. TextProtocol=Direct2026.01 | 58 | |
| Emb. Tokens (COCO)Protocol=CoT2026.01 | 58 | |
| SpaRE-7BModel size=7B2025.04 | 57.5 | |
| Qwen3.5Parameters=9B, Internal Reasoning (Think mode)=false2026.05 | 56.77 | |
| SpatialThinker2025.10 | 56.4 | |
| SpatialThinker-7B (Ours)Model Scale=7B, Training Dataset=STVQA-7K2025.11 | 56.4 | |
| Rotation TokensProtocol=CoT2026.01 | 56 | |
| SpatialReasoner-R12025.10 | 55.7 | |
| Qwen3VLParameters=30BA3B, Internal Reasoning (Think mode)=true2026.05 | 55.55 | |
| Qwen2.5-VL-7B + Vanilla GRPOModel Scale=7B, Training Protocol=Vanilla GRPO, Training Dataset=STVQA-7K2025.11 | 54.7 | |
| Mini-o3Mode=Zero-shot, Type=Visual Agents2026.05 | 54.5 | |
| Qwen3VLParameters=8B, Internal Reasoning (Think mode)=true2026.05 | 54.48 | |
| SpaRE-2BModel size=2B2025.04 | 54.4 | |
| Gemma4Parameters=26BA4B, Internal Reasoning (Think mode)=false2026.05 | 53.61 | |
| Qwen2.5-VL-7B + SFTModel Scale=7B, Training Protocol=SFT, Training Dataset=STVQA-7K2025.11 | 53.6 | |
| InternVL2-8BModel size=8B2025.04 | 53.3 | |
| SpatialThinker-3B (Ours)Model Scale=3B, Training Dataset=STVQA-7K2025.11 | 52.9 | |
| VLAA-Thinker-Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B, Model Type=Open-Source General2025.11 | 52.2 | |
| SpaceOmModel Type=Open-Source Spatial2025.11 | 52.2 | |
| DeepEyesMode=Zero-shot, Type=Visual Agents2026.05 | 51.6 | |
| LLaVA-NeXT-8BModel size=8B2025.04 | 51.1 | |
| SpaceThinkerModel Type=Open-Source Spatial2025.11 | 51.1 | |
| Qwen2.5-VL-3B + SFTModel Scale=3B, Training Protocol=SFT, Training Dataset=STVQA-7K2025.11 | 50.8 | |
| Qwen2.5-VL-3B + Vanilla GRPOModel Scale=3B, Training Protocol=Vanilla GRPO, Training Dataset=STVQA-7K2025.11 | 50.1 | |
| Qwen2VL-7BModel size=7B2025.04 | 49.2 | |
| Emb. Tokens (COCO)Protocol=Direct2026.01 | 49 | |
| Emb. Tokens (ViTPose)Protocol=Direct2026.01 | 49 | |
| Qwen2VL-2BModel size=2B2025.04 | 48.8 | |
| Qwen2.5-VL 7BType=Generalist Baseline2026.05 | 48.4 | |
| Qwen2.5-VL-7BModel Scale=7B, Model Type=Open-Source General2025.11 | 48.4 | |
| LLaVA-NeXT-8BModel Scale=8B, Model Type=Open-Source General2025.11 | 48.4 | |
| Spatial-RGPT-7B w/ depthModel Scale=7B, Model Type=Open-Source Spatial, Auxiliary Inputs=depth2025.11 | 48.4 | |
| Claude-3.5-Sonnet-0620Model Type=Proprietary2025.11 | 48.2 | |
| SATORI-R1Model Type=Open-Source Spatial2025.11 | 48 | |
| SpaceLLaVa2025.04 | 47.2 | |
| InternVL2-2BModel size=2B2025.04 | 46.7 | |
| GPT-4o2025.04 | 45.3 | |
| GPT-4o-0513Model Type=Proprietary2025.11 | 44.3 | |
| Qwen2.5-VL-3BModel Scale=3B, Model Type=Open-Source General2025.11 | 44 | |
| Cambrian-1-8BModel Scale=8B, Model Type=Open-Source General2025.11 | 42.2 | |
| SpaceLLaVA-13BModel Scale=13B, Model Type=Open-Source Spatial2025.11 | 42 | |
| SpatialBot-3BModel Scale=3B, Model Type=Open-Source Spatial2025.11 | 41.1 | |
| LLaVAModel Scale=1.5-13B2026.01 | 40 | |
| GPT-4o-mini2025.04 | 39.1 | |
| Random2025.04 | 20.9 |