Robot Manipulation on SimplerEnv WidowX Visual Matching
80.8Average Success RateInternVLA-A1.5
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| InternVLA-A1.52026.07 | 80.8 | 92.4 | 85.4 | 69.5 | 75.7 | |
| OA-WAMModel Category=World-Action Models2026.05 | 79.3 | 83 | 71.1 | 65 | 98.2 | |
| Xiaomi-Robotics-02026.07 | 79.2 | 95.8 | 62.5 | 75 | 83.3 | |
| PRTS2026.04 | 77.1 | 83.3 | 66.6 | 58.3 | 100 | |
| CoWVLAModel Category=World-Action Models2026.05 | 76 | 79.2 | 66.7 | 62.5 | 95.8 | |
| Embodied-R1.5-VLA2026.06 | 74 | 83.3 | 75 | 37.5 | 100 | |
| EO-12026.07 | 72.7 | 63.6 | 54.5 | 81.8 | 90.9 | |
| MemoryVLAModel Category=World-Action Models2026.05 | 71.9 | 75 | 75 | 37.5 | 100 | |
| InternVLA-M1Model Category=Vision-Language-Action Models2026.05 | 71.7 | 87.5 | 67.9 | 31.3 | 100 | |
| InternVLA-M12026.07 | 71.7 | 87.5 | 67.9 | 31.3 | 100 | |
| VITAModel Category=World-Action Models2026.05 | 71.5 | 84.2 | 68.8 | 37.5 | 95.6 | |
| NoTVLAInterface / Backbone=Qwen3-VL-4B + narrative sparse action2025.10 | 65.7 | — | — | — | — | |
| StarVLA-GR00TInterface / Backbone=Qwen3-VL-4B + dual-system action head2025.10 | 65.3 | — | — | — | — | |
| StarVLA-OFTInterface / Backbone=Qwen3-VL-4B + chunked continuous actions2025.10 | 64.6 | — | — | — | — | |
| GR00T-N1.52026.06 | 62 | 75.3 | 54.3 | 57 | 61.3 | |
| GR00T-N1.52026.04 | 61.9 | 75.3 | 54.3 | 57 | 61.3 | |
| GR00T-N1.5Interface / Backbone=dual-system foundation policy2025.10 | 61.9 | — | — | — | — | |
| GR00T-N1.52026.07 | 61.9 | 75.3 | 54.3 | 57 | 61.3 | |
| Qwen3-VL-PI2026.04 | 60.9 | 78.1 | 46.9 | 30.2 | 88.5 | |
| StarVLA-πInterface / Backbone=Qwen3-VL-4B + flow-matching action expert2025.10 | 60.9 | — | — | — | — | |
| F1-VLA2026.04 | 59.4 | 50 | 70.8 | 50 | 66.7 | |
| SoFarInterface / Backbone=language-grounded orientation2025.10 | 58.3 | — | — | — | — | |
| VLA-JEPAModel Category=World-Action Models2026.05 | 57.3 | 75 | 70.8 | 12.5 | 70.8 | |
| 𝜋0.52026.06 | 57.1 | 49.3 | 64.7 | 44.7 | 69.7 | |
| GR00T-N1.62026.06 | 57.1 | 64.5 | 65.5 | 5.5 | 93 | |
| π0.52026.07 | 57.1 | 49.3 | 64.7 | 44.7 | 69.7 | |
| CogACT2026.04 | 51.3 | 71.7 | 50.8 | 15 | 67.5 | |
| CogACTModel Category=Vision-Language-Action Models2026.05 | 51.3 | 71.7 | 50.8 | 15 | 67.5 | |
| CogACT2026.06 | 51.2 | 71.7 | 50.8 | 15 | 67.5 | |
| π0-Fast2026.04 | 48.3 | 29.1 | 21.9 | 10.8 | 66.6 | |
| π0-FASTInterface / Backbone=tokenized π action variant2025.10 | 48.3 | — | — | — | — | |
| UniVLAInterface / Backbone=task-centric latent actions2025.10 | 45.6 | — | — | — | — | |
| ThinkActModel Category=World-Action Models2026.05 | 43.8 | 58.3 | 37.5 | 8.7 | 70.8 | |
| SpatialVLA2026.04 | 42.7 | 16.7 | 25 | 29.2 | 100 | |
| SpatialVLAModel Category=Vision-Language-Action Models2026.05 | 42.7 | 16.7 | 25 | 29.2 | 100 | |
| SpatialVLAInterface / Backbone=PaliGemma2 + spatial tokens2025.10 | 42.7 | — | — | — | — | |
| SpatialVLA2026.06 | 42.7 | 16.7 | 25 | 29.2 | 100 | |
| OpenVLA-OFT2026.04 | 41.8 | 34.2 | 30 | 30 | 72.5 | |
| OpenVLA-OFT2026.06 | 41.8 | 34.2 | 30 | 30 | 72.5 | |
| 𝜋0-FAST2026.06 | 32.1 | 29.1 | 21.9 | 10.8 | 66.6 | |
| π0.52026.04 | 28.2 | 29.2 | 41.7 | 0 | 41.7 | |
| π02026.04 | 27.1 | 29.1 | 0 | 16.6 | 62.5 | |
| π0Interface / Backbone=VLA + flow expert2025.10 | 27.1 | — | — | — | — | |
| 𝜋02026.06 | 27.1 | 29.1 | 0 | 16.6 | 62.5 | |
| π02026.07 | 27.1 | 29.1 | 0 | 16.6 | 62.5 | |
| Octo-BaseInterface / Backbone=generalist robot policy2025.10 | 16 | — | — | — | — | |
| OpenVLAInterface / Backbone=open token VLA2025.10 | 4.2 | — | — | — | — | |
| OpenVLA2026.06 | 4.2 | 4.2 | 0 | 0 | 12.5 | |
| RT-1-X2026.04 | 1.1 | 0 | 4.2 | 0 | 0 | |
| RT-1-XInterface / Backbone=large-scale RT-style robot transformer2025.10 | 1.1 | — | — | — | — | |
| RT-1-X2026.06 | 1.1 | 0 | 4.2 | 0 | 0 | |
| OpenVLA2026.04 | 1 | 0 | 0 | 0 | 4.1 | |
| F1-VLAModel Category=World-Action Models2026.05 | — | 50 | 70.8 | 50 | 66.7 |