Goal-conditioned visual planning on COIN T=4 71 (test)
33.29Success Rate (SR)GeoWorld ViT-g384
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GeoWorld ViT-g384Model Category=Predictive (World) Models, Backbone=ViT-g3842026.02 | 33.29 | 61.56 | 91.84 | |
| GPT-5Model Category=LLM-Based2026.02 | 32.64 | 56.84 | 86.38 | |
| V-JEPA 2 ViT-g384Model Category=Predictive (World) Models, Backbone=ViT-g3842026.02 | 31.63 | 59.28 | 90.51 | |
| Gemini 2.5 ProModel Category=LLM-Based2026.02 | 30.2 | 60.13 | 84.82 | |
| GeoWorld ViT-gModel Category=Predictive (World) Models, Backbone=ViT-g2026.02 | 29.93 | 60.65 | 91.81 | |
| V-JEPA 2 ViT-gModel Category=Predictive (World) Models, Backbone=ViT-g2026.02 | 29.6 | 57.73 | 90.23 | |
| GeoWorld ViT-HModel Category=Predictive (World) Models, Backbone=ViT-H2026.02 | 28.82 | 58.1 | 90.48 | |
| V-JEPA 2 ViT-HModel Category=Predictive (World) Models, Backbone=ViT-H2026.02 | 27.38 | 56.07 | 89.2 | |
| GeoWorld ViT-LModel Category=Predictive (World) Models, Backbone=ViT-L2026.02 | 26.4 | 54.52 | 86.98 | |
| Qwen3-VL-MaxModel Category=LLM-Based2026.02 | 26.17 | 57.13 | 87.56 | |
| InternVL3.5-241BModel Category=LLM-Based2026.02 | 25.46 | 55.3 | 88.22 | |
| V-JEPA 2 ViT-LModel Category=Predictive (World) Models, Backbone=ViT-L2026.02 | 25.29 | 53.3 | 86.21 | |
| VideoWorldModel Category=Generative (World) Models2026.02 | 23.74 | 51.27 | 85.33 |