Visual Reasoning on SpatialQA Depth (test)
90Union ScoreChain-of-Thought
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Chain-of-ThoughtTraining Strategy=Supervised Post-Training2026.05 | 90 | 77.5 | |
| R1-Onevision-7BTraining Strategy=Previous Methods2026.05 | 87.5 | 72.5 | |
| Qwen2.5-VL-3BTraining Strategy=Previous Methods2026.05 | 86.3 | 73.8 | |
| VisionReasoner-7BTraining Strategy=Previous Methods2026.05 | 85 | 78.8 | |
| SFTTraining Strategy=Supervised Post-Training2026.05 | 85 | 65 | |
| DAPO+OursTraining Strategy=Reinforcement Post-Training2026.05 | 82.5 | 77.5 | |
| LLaVA-OV-7BTraining Strategy=Previous Methods2026.05 | 80 | 56.3 | |
| GRPOTraining Strategy=Reinforcement Post-Training2026.05 | 80 | 70 | |
| GRPO+OursTraining Strategy=Reinforcement Post-Training2026.05 | 80 | 73.8 | |
| DAPOTraining Strategy=Reinforcement Post-Training2026.05 | 76.3 | 73.8 |