Spatial Reasoning on MindCube
94.5AccuracyHuman
Evaluation Results
| Method | Links | |
|---|---|---|
| Human2026.02 | 94.5 | |
| HumanCategory=Baseline2026.02 | 94.5 | |
| SenseNova-SI (InternVL3-8B)Category=Open-Sourced spatial models2026.02 | 85.6 | |
| GeoThinker QWEN2.5VL-7BActive Perception=true2026.02 | 83.6 | |
| GeoThinker QWEN3VL-8BActive Perception=true2026.02 | 83 | |
| ACE-Brain-0-8BModel Type=Embodied Brain MLLM2026.03 | 82.1 | |
| GEMINI-3-PRO-PREVIEW2026.02 | 70.8 | |
| Gemini-3 ProCategory=Proprietary models2026.02 | 70.8 | |
| Grok-4Category=Proprietary models2026.02 | 63.5 | |
| SSR-2DCategory=Open-Sourced spatial models2026.02 | 63.5 | |
| Gemini 2.5 ProAccess Mode=API call only2026.05 | 59.2 | |
| GPT-5Access Mode=API call only2026.05 | 58 | |
| GEMINI-2.5-PRO2026.02 | 57.6 | |
| Gemini-2.5-ProData Scale=-, Training Phases=-2026.01 | 57.6 | |
| Gemini-2.5-ProModel Type=Closed-source MLLM2026.03 | 57.6 | |
| Molmo2-ERAccess Mode=Open weights, Open data, Open code2026.05 | 57 | |
| GPT-52026.02 | 56.3 | |
| GPT-5Category=Proprietary models2026.02 | 56.3 | |
| GPT-5-miniAccess Mode=API call only2026.05 | 55.6 | |
| GR-ER 1.5 ThinkingAccess Mode=API call only2026.05 | 54.7 | |
| GPT-4oData Scale=-, Training Phases=-2026.01 | 53.4 | |
| MindCube-3B-RawQA-SFTCategory=Open-Sourced spatial models2026.02 | 51.7 | |
| SEED-1.62026.02 | 48.7 | |
| Seed-1.6Category=Proprietary models2026.02 | 48.7 | |
| GR-ER 1.5Access Mode=API call only2026.05 | 47.7 | |
| LLaVA-OV-7BModel=LLaVA-OV-7B2026.05 | 47.3 | |
| GPT-4oModel Type=Closed-source MLLM2026.03 | 46.1 | |
| LLaVA-OV-7BAccess Mode=Open weights only2026.05 | 45.6 | |
| Qwen2.5-VL-7B + SAGEModel=Qwen2.5-VL-7B + SAGE2026.05 | 44.5 | |
| SPATIALLADDER-3B2026.02 | 43.4 | |
| SpatialLadder-3BData Scale=26K, Training Phases=3 (S+R)2026.01 | 43.4 | |
| SpatialLadder-3BModel=SpatialLadder-3B2026.05 | 43.4 | |
| InternVL3.5-4BAccess Mode=Open weights only2026.05 | 42.6 | |
| INTERNVL3-8B2026.02 | 41.5 | |
| InternVL3-8BData Scale=-, Training Phases=-2026.01 | 41.5 | |
| InternVL3-8BModel Type=Open-source general-purpose MLLM2026.03 | 41.5 | |
| VLM-3R-7BActive Perception=false2026.02 | 40 | |
| VLM-3R-7BCategory=Open-Sourced spatial models2026.02 | 40 | |
| VST-7B-SFT2026.02 | 39.7 | |
| VST-7B-SFTData Scale=4.1M, Training Phases=1 (S)2026.01 | 39.7 | |
| InternVL3.5-8BAccess Mode=Open weights only2026.05 | 39.7 | |
| CAMBRIAN-S-7B2026.02 | 39.6 | |
| SmoothOp-7BData Scale=50K, Training Phases=1 (R)2026.01 | 39.6 | |
| Cambrian-S-7BCategory=Open-Sourced spatial models2026.02 | 39.6 | |
| SmoothOp-3BData Scale=50K, Training Phases=1 (R)2026.01 | 39.4 | |
| GPT-4oModel=GPT-4o2026.05 | 38.8 | |
| SpaceR-7BData Scale=151K, Training Phases=2 (S+R)2026.01 | 37.9 | |
| QWEN2.5-VL-3B-INSTRUCT2026.02 | 37.6 | |
| Qwen2.5-VL-3BData Scale=-, Training Phases=-2026.01 | 37.6 | |
| Molmo2Access Mode=Open weights, Open data, Open code2026.05 | 37.6 | |
| INTERNVL3-2B2026.02 | 37.5 | |
| VG-LLM-4BActive Perception=false2026.02 | 36.9 | |
| Claude-4-SonnetModel Type=Closed-source MLLM2026.03 | 36.6 | |
| VG-LLM-8B*Active Perception=false, Training Setting=S1+S22026.02 | 36.1 | |
| QWEN2.5-VL-7B-INSTRUCT2026.02 | 36 | |
| Qwen2.5-VL-7BData Scale=-, Training Phases=-2026.01 | 36 | |
| Qwen2.5-VL-7B-Inst.Model Type=Open-source general-purpose MLLM2026.03 | 36 | |
| VST-3B-SFT2026.02 | 35.9 | |
| VST-3B-SFTData Scale=4.1M, Training Phases=1 (S)2026.01 | 35.9 | |
| ViLaSR-7BData Scale=81K, Training Phases=2 (S+R)2026.01 | 35.1 | |
| InternVL3.5-8BModel Type=Open-source general-purpose MLLM, Evaluation Protocol=Our evaluation framework2026.03 | 35.1 | |
| Ouro-Spatial-8BModel Parameters=8B2026.06 | 35.1 | |
| Qwen3-VL-8BModel Parameters=8B2026.06 | 35 | |
| BAGEL-7B-MOT2026.02 | 34.7 | |
| LLava-OneVision-7BData Scale=-, Training Phases=-2026.01 | 34.7 | |
| Bagel-7B-MoTCategory=Open-Sourced general models2026.02 | 34.7 | |
| Vlaser-8BModel Type=Embodied Brain MLLM, Evaluation Protocol=Our evaluation framework2026.03 | 34.6 | |
| Qwen3-VL-2B-Inst.Model Type=Open-source general-purpose MLLM2026.03 | 34.5 | |
| Qwen3-VL-8BAccess Mode=Open weights only2026.05 | 33.4 | |
| Ouro-Spatial-4BModel Parameters=4B2026.06 | 33.4 | |
| Random Choice2026.02 | 33 | |
| Random ChoiceCategory=Baseline2026.02 | 33 | |
| VG-LLM-8BActive Perception=false2026.02 | 32.7 | |
| CAMBRIAN-S-3B2026.02 | 32.5 | |
| MiMo-Embodied-7BModel Type=Embodied Brain MLLM, Evaluation Protocol=Our evaluation framework2026.03 | 32.3 | |
| VST-7B-SFTCategory=Open-Sourced spatial models2026.02 | 32 | |
| QWEN3-VL-2B-INSTRUCT2026.02 | 31.4 | |
| RoboBrain2.0-7BModel Type=Embodied Brain MLLM, Evaluation Protocol=Our evaluation framework2026.03 | 31.2 | |
| Pelican-VL-7BModel Type=Embodied Brain MLLM, Evaluation Protocol=Our evaluation framework2026.03 | 31 | |
| ViLaSR-7BCategory=Open-Sourced spatial models2026.02 | 30.2 | |
| VeBrain-7BModel Type=Embodied Brain MLLM, Evaluation Protocol=Our evaluation framework2026.03 | 30.1 | |
| QWEN3-VL-8B-INSTRUCT2026.02 | 29.8 | |
| Qwen2.5-VL-7BModel=Qwen2.5-VL-7B2026.05 | 29.5 | |
| Qwen3-VL-8B-Inst.Model Type=Open-source general-purpose MLLM2026.03 | 29.4 | |
| Qwen3-VL-4BModel Parameters=4B2026.06 | 28.4 | |
| RoboBrain2.5-8BModel Type=Embodied Brain MLLM, Evaluation Protocol=Our evaluation framework2026.03 | 28.1 | |
| SpatialLadder-3BCategory=Open-Sourced spatial models2026.02 | 27.4 | |
| SpaceR-7BCategory=Open-Sourced spatial models2026.02 | 27.4 | |
| Qwen3-VL-4BAccess Mode=Open weights only2026.05 | 27.4 | |
| Spatial-MLLM-4BCategory=Open-Sourced spatial models2026.02 | 26.1 | |
| InternVL-2.5-8BModel=InternVL-2.5-8B2026.05 | 18.7 |