Visual Planning on MAZE
74.5EMVPRL
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| VPRLInput=Image, Output=Image, post-trained=true, backbone=LVM-7B2025.05 | 74.5 | 77.6 | |
| Qwen 2.5-VL-Instruct-7B (SFT)Input=Text + Image, Output=Text, post-trained=true2025.05 | 60.9 | 70.3 | |
| VPFTInput=Image, Output=Image, post-trained=true, backbone=LVM-7B2025.05 | 59 | 64 | |
| Gemini 2.5 Pro (think)Input=Text + Image, Output=Text2025.05 | 21.5 | 35.5 | |
| Gemini 2.0 Flash (Direct)Input=Text + Image, Output=Text2025.05 | 8.3 | 31.4 | |
| Gemini 2.0 Flash (CoT)Input=Text + Image, Output=Text2025.05 | 6.9 | 29.8 | |
| Qwen 2.5-VL-Instruct-7B (CoT)Input=Text + Image, Output=Text2025.05 | 2.3 | 15.2 | |
| Qwen 2.5-VL-Instruct-7B (Direct)Input=Text + Image, Output=Text2025.05 | 0.6 | 14.5 |