Visual Question Answering on RoboSpatial-Home
78.1AccuracySpatialThinker-30B
Evaluation Results
| Method | Links | |
|---|---|---|
| SpatialThinker-30BTraining Dataset=STVQA-7K, Training Protocol=RL with Dense Rewards (Ours)2025.11 | 78.1 | |
| SpatialThinker-7BTraining Dataset=STVQA-7K, Training Protocol=RL with Dense Rewards (Ours)2025.11 | 76.3 | |
| Qwen2.5-VL-7B + Vanilla GRPOTraining Dataset=STVQA-7K, Training Protocol=Vanilla GRPO2025.11 | 76.2 | |
| Qwen2.5-VL-7B + SFTTraining Dataset=STVQA-7K, Training Protocol=SFT2025.11 | 72.4 | |
| GPT-5-0807Model Category=Proprietary and Open-Source MLLMs2025.11 | 71.5 | |
| Qwen2.5-VL-7BModel Category=Proprietary and Open-Source MLLMs2025.11 | 70.6 | |
| SpatialThinker-3BTraining Dataset=STVQA-7K, Training Protocol=RL with Dense Rewards (Ours)2025.11 | 70.6 | |
| Qwen2.5-VL-3B + SFTTraining Dataset=STVQA-7K, Training Protocol=SFT2025.11 | 69.8 | |
| Claude-4-Sonnet-0514Model Category=Proprietary and Open-Source MLLMs2025.11 | 69.7 | |
| VLAA-Thinker-7BModel Category=Proprietary and Open-Source MLLMs2025.11 | 68.9 | |
| SpaceOmModel Category=Proprietary and Open-Source MLLMs2025.11 | 68.9 | |
| GPT-4o-0513Model Category=Proprietary and Open-Source MLLMs2025.11 | 68.4 | |
| Qwen2.5-VL-3B + Vanilla GRPOTraining Dataset=STVQA-7K, Training Protocol=Vanilla GRPO2025.11 | 64 | |
| Qwen2.5-VL-3BModel Category=Proprietary and Open-Source MLLMs2025.11 | 58.7 | |
| Claude-3.5-Sonnet-0620Model Category=Proprietary and Open-Source MLLMs2025.11 | 57 | |
| Qwen3-VL-30BModel Category=Proprietary and Open-Source MLLMs2025.11 | 53.1 | |
| SpaceThinkerModel Category=Proprietary and Open-Source MLLMs2025.11 | 52.6 |