Spatial Reasoning on SAT Real
75.3Accuracy (Pass@1)DCRL
Evaluation Results
| Method | Links | |
|---|---|---|
| DCRLModel Category=Our Models2026.06 | 75.3 | |
| World2VLM-GRPOSetting=SVC as WM2026.04 | 72.67 | |
| Qwen3-VL-8BModel Category=Our Models2026.06 | 70 | |
| World2VLM-GRPOSetting=HY-WorldPlay as WM2026.04 | 69.33 | |
| GPT-4o (CoT)Reasoning Strategy=CoT2026.04 | 68.67 | |
| World2VLM-SFTSetting=HY-WorldPlay as WM2026.04 | 68.66 | |
| GRPO-TBackbone=Qwen2.5-VL-7B-Instruct, Model Scale=7B, Reasoning Strategy=RL (Task Reward Only)2026.04 | 67 | |
| FGRPOBackbone=Qwen2.5-VL-7B-Instruct, Model Scale=7B, Reasoning Strategy=Faithful GRPO2026.04 | 65.66 | |
| VL-RethinkerReasoning Strategy=MRM baseline2026.04 | 65 | |
| GPT-5-nano (CoT)Reasoning Strategy=CoT2026.04 | 64 | |
| World2VLM-SFTSetting=SVC as WM2026.04 | 64 | |
| Non-reasoningBackbone=Qwen2.5-VL-7B-Instruct, Model Scale=7B, Reasoning Strategy=Non-reasoning2026.04 | 63.11 | |
| TreeVGRReasoning Strategy=MRM baseline2026.04 | 61 | |
| CoT promptingBackbone=Qwen2.5-VL-7B-Instruct, Model Scale=7B, Reasoning Strategy=CoT prompting2026.04 | 59.22 | |
| Non-reasoningBackbone=Qwen2.5-VL-3B-Instruct, Model Scale=3B, Reasoning Strategy=Non-reasoning2026.04 | 59 | |
| GRPO-TBackbone=Qwen2.5-VL-3B-Instruct, Model Scale=3B, Reasoning Strategy=RL (Task Reward Only)2026.04 | 58.67 | |
| FGRPOBackbone=Qwen2.5-VL-3B-Instruct, Model Scale=3B, Reasoning Strategy=Faithful GRPO2026.04 | 58.6 | |
| Vision-R1Reasoning Strategy=MRM baseline2026.04 | 58.45 | |
| ViGoRL-SpatialReasoning Strategy=MRM baseline2026.04 | 58.44 | |
| GPT-4oModel Category=Proprietary Model2026.06 | 57.5 | |
| Qwen2.5-VL-7BModel Category=Open-Weight Multi Image Models2026.06 | 56.33 | |
| CoT promptingBackbone=Qwen2.5-VL-3B-Instruct, Model Scale=3B, Reasoning Strategy=CoT prompting2026.04 | 55 | |
| R1-OneVisionReasoning Strategy=MRM baseline2026.04 | 51.5 | |
| Qwen2.5-VL-7BSetting=Base VLM2026.04 | 44.67 | |
| Qwen2.5-VL-7B + WMSetting=MindJourney-style[55]2026.04 | 31.33 |