Robotic Manipulation on PrimitiveSkill
100Place Success RateQwen2.5-VL-72B
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Qwen2.5-VL-72BBackbone=Qwen2.5-VL-72B2026.05 | 100 | 50 | 0 | 100 | 44 | |
| VAGEN-FullBackbone=Qwen2.5-VL-3B, Reward Strategy=Dense Extrinsic, Reasoning Strategy=World Model Reasoning for Visual States2026.05 | 100 | 88 | 100 | 100 | 97 | |
| GLANCE-FullBackbone=Qwen2.5-VL-3B, Reward Strategy=Dense Extrinsic, Reasoning Strategy=World Model Reasoning for Visual States2026.05 | 100 | 88 | 100 | 100 | 97 | |
| VAGEN-BaseBackbone=Qwen2.5-VL-3B, Reward Strategy=Sparse Extrinsic, Reasoning Strategy=World Model Reasoning Strategy2026.05 | 100 | 88 | 88 | 88 | 91 | |
| GLANCE-BaseBackbone=Qwen2.5-VL-3B, Reward Strategy=Sparse Extrinsic, Reasoning Strategy=World Model Reasoning Strategy2026.05 | 100 | 88 | 100 | 88 | 94 | |
| GLANCE w/ Turn-PPOBackbone=Qwen2.5-VL-3B, Reward Strategy=Turn-level PPO, Reasoning Strategy=World Model Reasoning Strategy2026.05 | 100 | 63 | 0 | 100 | 66 | |
| o4-mini2026.05 | 100 | 50 | 0 | 75 | 56 | |
| Gemini 2.5 Pro2026.05 | 63 | 63 | 0 | 75 | 50 | |
| Claude 4.5 Sonnet2026.05 | 63 | 50 | 0 | 100 | 53 | |
| Claude 3.7 Sonnet2026.05 | 63 | 13 | 0 | 100 | 44 | |
| GPT-4o2026.05 | 50 | 63 | 0 | 88 | 50 | |
| Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B2026.05 | 0 | 0 | 0 | 75 | 19 | |
| Qwen2.5-VL-3BBackbone=Qwen2.5-VL-3B2026.05 | 0 | 0 | 0 | 0 | 0 | |
| VLM-R1-3BBackbone=VLM-R1-3B2026.05 | 0 | 0 | 0 | 0 | 0 | |
| Turn-PPO w/ MaskBackbone=Qwen2.5-VL-3B, Reward Strategy=Turn-level PPO, Reasoning Strategy=World Model Reasoning Strategy2026.05 | 0 | 0 | 0 | 100 | 25 | |
| Vanilla-PPOBackbone=Qwen2.5-VL-3B, Reward Strategy=RL Baseline, Reasoning Strategy=World Model Reasoning Strategy2026.05 | 0 | 0 | 0 | 0 | 0 | |
| GRPO w/ MaskBackbone=Qwen2.5-VL-3B, Reward Strategy=RL Baseline, Reasoning Strategy=World Model Reasoning Strategy2026.05 | 0 | 0 | 0 | 100 | 25 |