Visual Agentic Reasoning on Sokoban
85Success RateGLANCE-Full
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GLANCE-FullBackbone=Qwen2.5-VL-3B, Reward Strategy=Dense Extrinsic, Reasoning Strategy=World Model Reasoning for Visual States2026.05 | 85 | — | |
| VAGEN-FullBackbone=Qwen2.5-VL-3B, Reward Strategy=Dense Extrinsic, Reasoning Strategy=World Model Reasoning for Visual States2026.05 | 79 | — | |
| GLANCE-BaseBackbone=Qwen2.5-VL-3B, Reward Strategy=Sparse Extrinsic, Reasoning Strategy=World Model Reasoning Strategy2026.05 | 74 | — | |
| RewardFlowType=RL Training, Backbone=Qwen2.5-(VL)-7B-Instruct2026.03 | 62.4 | 3 | |
| VAGEN-BaseBackbone=Qwen2.5-VL-3B, Reward Strategy=Sparse Extrinsic, Reasoning Strategy=World Model Reasoning Strategy2026.05 | 61 | — | |
| Gemini 2.5 Pro2026.05 | 58 | — | |
| GLANCE w/ Turn-PPOBackbone=Qwen2.5-VL-3B, Reward Strategy=Turn-level PPO, Reasoning Strategy=World Model Reasoning Strategy2026.05 | 52 | — | |
| RewardFlowType=RL Training, Backbone=Qwen2.5-(VL)-3B-Instruct2026.03 | 49.2 | 2.2 | |
| o4-mini2026.05 | 44 | — | |
| GPT-4o2026.05 | 43 | — | |
| Turn-PPO w/ MaskBackbone=Qwen2.5-VL-3B, Reward Strategy=Turn-level PPO, Reasoning Strategy=World Model Reasoning Strategy2026.05 | 38 | — | |
| GiGPOType=RL Training, Backbone=Qwen2.5-(VL)-7B-Instruct2026.03 | 34.4 | 1.4 | |
| Claude 4.5 Sonnet2026.05 | 31 | — | |
| GRPOType=RL Training, Backbone=Qwen2.5-(VL)-3B-Instruct2026.03 | 26.6 | 1.3 | |
| Claude 3.7 Sonnet2026.05 | 25 | — | |
| GRPOType=RL Training, Backbone=Qwen2.5-(VL)-7B-Instruct2026.03 | 23.4 | 1 | |
| RLOOType=RL Training, Backbone=Qwen2.5-(VL)-3B-Instruct2026.03 | 22.7 | 1 | |
| GiGPOType=RL Training, Backbone=Qwen2.5-(VL)-3B-Instruct2026.03 | 21.9 | 1.2 | |
| RLOOType=RL Training, Backbone=Qwen2.5-(VL)-7B-Instruct2026.03 | 21.9 | 1 | |
| Qwen2.5-VL-72BBackbone=Qwen2.5-VL-72B2026.05 | 20 | — | |
| GRPO w/ MaskBackbone=Qwen2.5-VL-3B, Reward Strategy=RL Baseline, Reasoning Strategy=World Model Reasoning Strategy2026.05 | 20 | — | |
| BaseType=Prompting, Backbone=Qwen2.5-(VL)-7B-Instruct2026.03 | 18.8 | 0.9 | |
| Vanilla-PPOBackbone=Qwen2.5-VL-3B, Reward Strategy=RL Baseline, Reasoning Strategy=World Model Reasoning Strategy2026.05 | 18 | — | |
| VLM-R1-3BBackbone=VLM-R1-3B2026.05 | 16 | — | |
| BaseType=Prompting, Backbone=Qwen2.5-(VL)-3B-Instruct2026.03 | 14.1 | 0.5 | |
| Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B2026.05 | 14 | — | |
| Qwen2.5-VL-3BBackbone=Qwen2.5-VL-3B2026.05 | 13 | — |