Visual Spatial Planning on VSP (test)
99Average AccuracyLatentUMVis-Plan
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| LatentUMVis-PlanPlanning Paradigm=fine2026.04 | 99 | 100 | 100 | 100 | 97 | |
| LatentUMVis-PlanPlanning Paradigm=coarse2026.04 | 85 | 100 | 85 | 83 | 71 | |
| UniCanvasInput=Image, Text, Output=Image, (Text)2026.06 | 77 | 92 | 81 | 79 | 57 | |
| Mirage2026.04 | 76 | 93 | 83 | 76 | 51 | |
| ThinkMorph2026.04 | 76 | — | — | — | — | |
| Qwen2.5-VL-3BInput=Image, Text, Output=Text2026.06 | 72 | 88 | 82 | 73 | 47 | |
| BAGEL-TextInput=Image, Text, Output=Text2026.06 | 55 | 72 | 60 | 52 | 34 | |
| GPT-4o Zero-ShotInput=Image, Text, Output=Text2026.06 | 46 | 68 | 58 | 35 | 24 | |
| BAGEL-InterleaveInput=Image, Text, Output=Image, Text2026.06 | 24 | 32 | 23 | 26 | 13 | |
| InternVL3.5Model Scale=38B2026.04 | 20 | — | — | — | — | |
| MVoTInput=Image, Text, Output=Image, Text2026.06 | 12 | 24 | 13 | 9 | 3 | |
| MVoT2026.04 | 11 | 21 | 11 | 8 | 3 | |
| InternVL3.5Model Scale=8B2026.04 | 8 | — | — | — | — | |
| MMaDAInput=Image, Text, Output=Image, Text2026.06 | 8 | 18 | 9 | 2 | 2 | |
| AnoleInput=Image, Text, Output=Image, Text2026.06 | 6 | 13 | 8 | 3 | 1 | |
| Anole2026.04 | 1 | 2 | 1 | 0 | 0 | |
| Chameleon2026.04 | 1 | — | — | — | — |