Visual Question Answering on HallusionBench
75.2Simple AccuracySpatialThinker-30B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| SpatialThinker-30BTraining Dataset=STVQA-7K, Training Protocol=RL with Dense Rewards (Ours)2025.11 | 75.2 | — | — | |
| GPT-5-0807Model Category=Proprietary and Open-Source MLLMs2025.11 | 73.8 | — | — | |
| Claude-4-Sonnet-0514Model Category=Proprietary and Open-Source MLLMs2025.11 | 71.2 | — | — | |
| VLAA-Thinker-7BModel Category=Proprietary and Open-Source MLLMs2025.11 | 68.9 | — | — | |
| SpatialThinker-7BTraining Dataset=STVQA-7K, Training Protocol=RL with Dense Rewards (Ours)2025.11 | 66.4 | — | — | |
| Qwen2.5-VL-7B + SFTTraining Dataset=STVQA-7K, Training Protocol=SFT2025.11 | 66.2 | — | — | |
| SpaceThinkerModel Category=Proprietary and Open-Source MLLMs2025.11 | 65.4 | — | — | |
| SpaceOmModel Category=Proprietary and Open-Source MLLMs2025.11 | 62.9 | — | — | |
| SpatialThinker-3BTraining Dataset=STVQA-7K, Training Protocol=RL with Dense Rewards (Ours)2025.11 | 62.5 | — | — | |
| Qwen3-VL-30BModel Category=Proprietary and Open-Source MLLMs2025.11 | 61.5 | — | — | |
| Qwen2.5-VL-7B + Vanilla GRPOTraining Dataset=STVQA-7K, Training Protocol=Vanilla GRPO2025.11 | 60.7 | — | — | |
| Qwen2.5-VL-3B + Vanilla GRPOTraining Dataset=STVQA-7K, Training Protocol=Vanilla GRPO2025.11 | 59 | — | — | |
| Qwen2.5-VL-3B + SFTTraining Dataset=STVQA-7K, Training Protocol=SFT2025.11 | 58.9 | — | — | |
| Claude-3.5-Sonnet-0620Model Category=Proprietary and Open-Source MLLMs2025.11 | 55.5 | — | — | |
| GPT-4o-0513Model Category=Proprietary and Open-Source MLLMs2025.11 | 55 | — | — | |
| Qwen2.5-VL-7BModel Category=Proprietary and Open-Source MLLMs2025.11 | 52.9 | — | — | |
| LLaVA-1.5 13BIntervention Strategy=Q4 system redistribution (prop)2026.01 | 51.31 | 15.58 | 65.62 | |
| LLaVA-1.5 13BIntervention Strategy=No-intervention baseline2026.01 | 50.89 | 15.81 | 67.93 | |
| Qwen2.5-VL-3BModel Category=Proprietary and Open-Source MLLMs2025.11 | 46.3 | — | — |