Goal-conditioned visual planning on CrossTask T=3 88 (test)
51.71Success Rate (SR)GeoWorld ViT-g384
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GeoWorld ViT-g384Model Category=Predictive (World) Models, Backbone=ViT-g3842026.02 | 51.71 | 77.3 | 92.95 | |
| V-JEPA 2 ViT-g384Model Category=Predictive (World) Models, Backbone=ViT-g3842026.02 | 50.16 | 74.86 | 91.73 | |
| GPT-5Model Category=LLM-Based2026.02 | 50.03 | 72.38 | 91.18 | |
| GeoWorld ViT-gModel Category=Predictive (World) Models, Backbone=ViT-g2026.02 | 49.23 | 76.64 | 90.61 | |
| Gemini 2.5 ProModel Category=LLM-Based2026.02 | 48.91 | 73.82 | 90.3 | |
| V-JEPA 2 ViT-gModel Category=Predictive (World) Models, Backbone=ViT-g2026.02 | 48.13 | 73.42 | 89.62 | |
| GeoWorld ViT-HModel Category=Predictive (World) Models, Backbone=ViT-H2026.02 | 47.79 | 74.42 | 88.84 | |
| V-JEPA 2 ViT-HModel Category=Predictive (World) Models, Backbone=ViT-H2026.02 | 46.02 | 71.98 | 87.29 | |
| Qwen3-VL-MaxModel Category=LLM-Based2026.02 | 45.47 | 70.93 | 86.18 | |
| GeoWorld ViT-LModel Category=Predictive (World) Models, Backbone=ViT-L2026.02 | 44.8 | 70.54 | 86.3 | |
| InternVL3.5-241BModel Category=LLM-Based2026.02 | 44.03 | 70.01 | 84.41 | |
| V-JEPA 2 ViT-LModel Category=Predictive (World) Models, Backbone=ViT-L2026.02 | 43.36 | 69.55 | 84.75 | |
| VideoWorldModel Category=Generative (World) Models2026.02 | 41.59 | 66.11 | 82.64 |