Multi-view spatial reasoning on MindCube
77AccuracyViewFusion
Evaluation Results
| Method | Links | |
|---|---|---|
| ViewFusionSize=4B, Training Protocol=SFT + RL2026.03 | 77 | |
| SpatialClawBackbone=Gemma4-31B2026.06 | 72.8 | |
| Gemini-3-Pro-PreviewSize=-2026.03 | 70.8 | |
| ViewFusionSize=4B, Training Protocol=SFT2026.03 | 68.5 | |
| pySpatialBackbone=Gemma4-31B2026.06 | 67.1 | |
| Gemini-2.5-ProSize=-2026.03 | 57.6 | |
| GPT-5Size=-2026.03 | 56.3 | |
| SpaceTools ToolshedBackbone=Gemma4-31B2026.06 | 52.9 | |
| SpatialLadder-3BSize=3B2026.03 | 43.4 | |
| InternVL3-8BSize=8B2026.03 | 41.5 | |
| No-ThinkVDrop=✗2026.05 | 41.1 | |
| Cambrian-S-7BSize=7B2026.03 | 39.6 | |
| ThinkMorph (24K)VDrop=✗2026.05 | 39.2 | |
| VST-7B-RLSize=7B, Training Protocol=RL2026.03 | 39.1 | |
| SpaceR-7BSize=7B2026.03 | 37.9 | |
| Qwen2.5-VL-3B-InstructSize=3B, Training Protocol=Instruct2026.03 | 37.6 | |
| InternVL3-2BSize=2B2026.03 | 37.5 | |
| Qwen3-VL-4B-InstructSize=4B, Training Protocol=Instruct2026.03 | 37 | |
| PanoramicVDrop=✗2026.05 | 36.9 | |
| Top-down + VDropVDrop=✓2026.05 | 36.5 | |
| VST-3B-RLSize=3B, Training Protocol=RL2026.03 | 36.4 | |
| Qwen2.5-VL-7B-InstructSize=7B, Training Protocol=Instruct2026.03 | 36 | |
| Point MatchingVDrop=✗2026.05 | 35.2 | |
| Top-downVDrop=✗2026.05 | 35.2 | |
| ViLaSR-7BSize=7B2026.03 | 35.1 | |
| Qwen3-VL-2B-InstructSize=2B, Training Protocol=Instruct2026.03 | 34.5 | |
| Qwen3-VL-8BVDrop=✗2026.05 | 34.4 | |
| Panoramic + VDropVDrop=✓2026.05 | 34.1 | |
| Point Matching + VDropVDrop=✓2026.05 | 34.1 | |
| Spatial-MLLM-4BSize=4B2026.03 | 33.4 | |
| RandomChoiceSize=-2026.03 | 33 | |
| Cambrian-S-3BSize=3B2026.03 | 32.5 | |
| BAGELVDrop=✗2026.05 | 31.7 | |
| Qwen3-VL-8B-InstructSize=8B, Training Protocol=Instruct2026.03 | 29.4 | |
| Qwen3-VL-4BVDrop=✗2026.05 | 29 | |
| Text CoTVDrop=✗2026.05 | 25.3 | |
| BAGEL-Zebra-CoT (182K)VDrop=✗2026.05 | 21.7 |