Robot Manipulation on LIBERO (Spatial, Object, Goal, Long, Avg Metrics and Ranks)
99Spatial Successω-EVA Stage 3
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| ω-EVA Stage 3Core Params=1.2B, Robot Pretrain=No2026.06 | 99 | 99.8 | 98.2 | 97.4 | 98.6 | — | — | — | |
| π0.5Type=VLA, Params=3B, EPT=w/2026.06 | 98.8 | 98.2 | 98 | 92.4 | 96.9 | — | 6 | — | |
| π0.5Core Params=3.3B, Robot Pretrain=Yes2026.06 | 98.8 | 98.2 | 98 | 92.4 | 96.9 | — | — | — | |
| ω-EVA Stage 2Core Params=0.8B, Robot Pretrain=No2026.06 | 98.8 | 99.4 | 97.6 | 95.8 | 97.9 | — | — | — | |
| LingBot-VAType=WAM, Params=5.3B, EPT=w/2026.06 | 98.5 | 99.6 | 97.2 | 98.5 | 98.5 | — | 1 | — | |
| LingBot-VACore Params=5.3B, Robot Pretrain=Yes2026.06 | 98.5 | 99.6 | 97.2 | 98.5 | 98.5 | — | — | — | |
| GeoVLAView Setting=Multi-view VLA, Category=3D-cloud-enhanced VLA2025.10 | 98.4 | 99 | 96.6 | 96.6 | 97.7 | — | — | — | |
| LightVLA (OpenVLA-OFT)Model Size=7B, Latency (ms)=91.562025.06 | 98.4 | 98.4 | 98.2 | 94.6 | 97.4 | — | — | — | |
| Light-WAMType=WAM, Params=2B, EPT=w/o2026.06 | 98.2 | 99.6 | 97.8 | 93 | 97.2 | 1 | 3 | — | |
| Fast-WAMCore Params=6B, Robot Pretrain=No2026.06 | 98.2 | 100 | 97 | 95.2 | 97.6 | — | — | — | |
| Fast-WAM2026.06 | 98.2 | 100 | 97 | 95.2 | 97.6 | — | — | — | |
| HiMem-WAM2026.06 | 98.2 | 99.8 | 98.4 | 94.5 | 97.7 | — | — | — | |
| 3D-CAVLAView Setting=Multi-view VLA, Category=Depth-enhanced VLA2025.10 | 98.2 | 99.8 | 98.2 | 96.1 | 98.1 | — | — | — | |
| GR00T-RLRCModel Size=2B, Latency (ms)=55.42025.06 | 98.2 | 98 | 97.8 | 93.4 | 96.9 | — | — | — | |
| GR00T N1.6Model Size=3B, Latency (ms)=752025.06 | 98 | 97.8 | 98.4 | 95.4 | 97.4 | — | — | — | |
| OpenVLA-OFT-RLRCModel Size=2B, Latency (ms)=65.592025.06 | 97.8 | 99.6 | 98.2 | 94.8 | 97.6 | — | — | — | |
| OpenVLA-OFTType=VLA, Params=7B, EPT=w/2026.06 | 97.6 | 98.4 | 97.9 | 94.5 | 97.1 | — | 4 | — | |
| QDepth-VLAView Setting=Multi-view VLA, Category=Depth-enhanced VLA2025.10 | 97.6 | 96.6 | 95.2 | 90 | 94.9 | — | — | — | |
| OpenVLA-OFTModel Size=7B, Latency (ms)=149.232025.06 | 97.6 | 98.4 | 97.9 | 94.5 | 97.1 | — | — | — | |
| DreamVLAView Setting=Multi-view VLA, Category=Depth-enhanced VLA2025.10 | 97.5 | 94 | 89.5 | 89.5 | 92.6 | — | — | — | |
| villa-XPretraining Method=Mix Videos2025.09 | 97.5 | 97 | 91.5 | 74.5 | 90.1 | — | — | — | |
| NORA-1.52026.06 | 97.3 | 96.4 | 94.5 | 89.6 | 94.5 | — | — | — | |
| SwiftVLAHistory=true, 3D=true2026.06 | 97.2 | 96.8 | 97.4 | 89 | 95.1 | — | — | — | |
| Fast-WAMType=WAM, Params=6B, EPT=w/o2026.06 | 97 | 99.4 | 96.6 | 94.8 | 97 | 2 | 5 | — | |
| π0Type=VLA, Params=3B, EPT=w/2026.06 | 96.8 | 98.8 | 95.8 | 85.2 | 94.1 | — | 8 | — | |
| MotusType=WAM, Params=8B, EPT=w/2026.06 | 96.8 | 99.8 | 96.6 | 97.6 | 97.7 | — | 2 | — | |
| π0History=false, 3D=false2026.06 | 96.8 | 98.8 | 95.8 | 85.2 | 94.2 | — | — | — | |
| π0Core Params=3.3B, Robot Pretrain=Yes2026.06 | 96.8 | 98.8 | 95.8 | 85.2 | 94.1 | — | — | — | |
| MotusCore Params=8B, Robot Pretrain=Yes2026.06 | 96.8 | 99.8 | 96.6 | 97.6 | 97.7 | — | — | — | |
| π02026.06 | 96.8 | 98.8 | 95.8 | 85.2 | 94.2 | — | — | — | |
| π0 finetunedView Setting=Multi-view VLA, Category=General VLA2025.10 | 96.8 | 98.8 | 95.8 | 85.2 | 94.2 | — | — | — | |
| π0Model Size=3B2025.06 | 96.8 | 98.8 | 95.8 | 85.2 | 94.1 | — | — | — | |
| UniVLAView Setting=Multi-view VLA, Category=General VLA2025.10 | 96.5 | 96.8 | 95.6 | 92 | 95.2 | — | — | — | |
| EfficientVLA (OpenVLA-OFT)Model Size=7B, Latency (ms)=99.482025.06 | 96.5 | 91.1 | 96 | 72.1 | 88.9 | — | — | — | |
| π0-FASTHistory=false, 3D=false2026.06 | 96.4 | 96.8 | 88.6 | 60.2 | 85.5 | — | — | — | |
| DepthVLAHistory=false, 3D=true2026.06 | 96.4 | 98 | 95.8 | 89.2 | 94.9 | — | — | — | |
| AtomVLA2026.06 | 96.4 | 99.6 | 97.6 | 94.4 | 97 | — | — | — | |
| π0-FAST finetunedView Setting=Multi-view VLA, Category=General VLA2025.10 | 96.4 | 96.8 | 88.6 | 60.2 | 85.5 | — | — | — | |
| π0-FastModel Size=3B2025.06 | 96.4 | 96.8 | 88.6 | 60.2 | 85.5 | — | — | — | |
| MotionVLAHistory=true, 3D=true2026.06 | 96.2 | 98 | 96.2 | 91.2 | 95.4 | — | — | — | |
| VLA-JEPACore Params=3B, Robot Pretrain=Yes2026.06 | 96.2 | 99.6 | 97.2 | 95.8 | 97.2 | — | — | — | |
| VLA-AdapterType=VLA, Params=0.6B, EPT=w/o2026.06 | 96 | 96.8 | 97.4 | 94.4 | 96.2 | 3 | 7 | — | |
| HiMem-WAM (w/o Stage II)Ablation=Without Stage II Latent Action Pretraining2026.06 | 96 | 99.6 | 97.1 | 93.8 | 96.6 | — | — | — | |
| OriginalModel=X-VLA, Compression rate=0.12026.06 | 95 | 98.5 | 93.5 | 87.5 | 93.63 | — | — | 91.72 | |
| VLA-JEPACore Params=3B, Robot Pretrain=No2026.06 | 94.8 | 99.6 | 95.8 | 94 | 96.1 | — | — | — | |
| GR00T-N1History=false, 3D=false2026.06 | 94.4 | 97.6 | 93 | 90.6 | 93.9 | — | — | — | |
| EinSortModel=X-VLA, Compression rate=0.12026.06 | 93.5 | 97.5 | 90.5 | 88 | 92.38 | — | — | 90.33 | |
| SmolVLA (SmolVLM-2.25B)Pretraining Method=VLM Checkpoint2025.09 | 93 | 94 | 91 | 77 | 88.75 | — | — | — | |
| SmolVLAModel Size=2.25B2025.06 | 93 | 94 | 91 | 77 | 88.8 | — | — | — | |
| Diffusion Policy w/ latent (Ours)Pretraining Method=Human Videos2025.09 | 92.7 | 96.3 | 94.2 | 88 | 92.8 | — | — | — | |
| UniVLAPretraining Method=Human Videos2025.09 | 91.2 | 94.2 | 90.2 | 79.4 | 88.7 | — | — | — | |
| π0Pretraining Method=Robot Actions2025.09 | 90 | 86 | 95 | 73 | 86 | — | — | — | |
| 4D-VLAHistory=true, 3D=true2026.06 | 88.9 | 95.2 | 90.9 | 79.1 | 88.6 | — | — | — | |
| 4D-VLAView Setting=Multi-view VLA, Category=Depth-enhanced VLA2025.10 | 88.9 | 95.2 | 90.9 | 79.1 | 88.6 | — | — | — | |
| PixelVLAHistory=false, 3D=false2026.06 | 88.5 | 90 | 85.8 | 82.6 | 86.7 | — | — | — | |
| SpatialVLAHistory=false, 3D=true2026.06 | 88.2 | 89.9 | 78.6 | 55.5 | 78.1 | — | — | — | |
| SpatialVLA2026.06 | 88.2 | 89.9 | 78.6 | 55.5 | 78.1 | — | — | — | |
| SpatialVLAView Setting=Single-view VLA, Category=3D-cloud-enhanced VLA2025.10 | 88.2 | 89.9 | 78.6 | 55.5 | 78.1 | — | — | — | |
| SpatialVLAModel Size=7B2025.06 | 88.2 | 89.9 | 78.6 | 55.5 | 78.1 | — | — | — | |
| Diffusion Policy w/o latentPretraining Method=None2025.09 | 88.1 | 81.4 | 87.5 | 70.5 | 81.87 | — | — | — | |
| WorldVLA2026.06 | 87.6 | 96.2 | 83.4 | 60 | 81.8 | — | — | — | |
| CoT-VLAHistory=false, 3D=false2026.06 | 87.5 | 91.6 | 87.6 | 69 | 83.9 | — | — | — | |
| CoT-VLA-7BView Setting=Single-view VLA, Category=General VLA2025.10 | 87.5 | 91.6 | 87.6 | 69 | 81.1 | — | — | — | |
| π0 (Paligemma-3B)Pretraining Method=VLM Checkpoint2025.09 | 87 | 63 | 89 | 48 | 71.8 | — | — | — | |
| VLA-OSModel Size=0.5B2025.06 | 87 | 96.5 | 92.7 | 66 | 85.6 | — | — | — | |
| 3D-CAVLAView Setting=Single-view VLA, Category=Depth-enhanced VLA2025.10 | 86.1 | 94.7 | 82.9 | 66.8 | 82.6 | — | — | — | |
| QDepth-VLAView Setting=Single-view VLA, Category=Depth-enhanced VLA2025.10 | 86 | 88.8 | 94 | 72.6 | 85.4 | — | — | — | |
| OpenVLA-RLRCModel Size=2B, Latency (ms)=74.042025.06 | 85.2 | 88 | 79.6 | 52.8 | 76.4 | — | — | — | |
| OpenVLAType=VLA, Params=7B, EPT=w/2026.06 | 84.7 | 88.4 | 79.2 | 53.7 | 76.5 | — | 9 | — | |
| OpenVLAHistory=false, 3D=false2026.06 | 84.7 | 88.4 | 79.2 | 53.7 | 76.5 | — | — | — | |
| OpenVLACore Params=7B, Robot Pretrain=Yes2026.06 | 84.7 | 88.4 | 79.2 | 53.7 | 76.5 | — | — | — | |
| OpenVLA2026.06 | 84.7 | 88.4 | 79.2 | 53.7 | 76.5 | — | — | — | |
| OpenVLA finetunedView Setting=Single-view VLA, Category=General VLA2025.10 | 84.7 | 88.4 | 79.2 | 53.7 | 76.5 | — | — | — | |
| OpenVLAPretraining Method=Robot Actions2025.09 | 84.7 | 88.4 | 79.2 | 53.7 | 76.5 | — | — | — | |
| OpenVLAModel Size=7B, Latency (ms)=1692025.06 | 84.7 | 88.4 | 79.2 | 53.7 | 76.5 | — | — | — | |
| TraceVLAHistory=true, 3D=false2026.06 | 84.6 | 85.2 | 75.1 | 54.1 | 74.8 | — | — | — | |
| VLA-Cache (OpenVLA)Model Size=7B, Latency (ms)=125.182025.06 | 83.8 | 85.8 | 76.4 | 52.8 | 74.7 | — | — | — | |
| OctoHistory=false, 3D=false2026.06 | 78.9 | 85.7 | 84.6 | 51.1 | 75.1 | — | — | — | |
| Octo finetunedView Setting=Multi-view VLA, Category=General VLA2025.10 | 78.9 | 85.7 | 84.6 | 51.1 | 75.1 | — | — | — | |
| OctoPretraining Method=Robot Actions2025.09 | 78.9 | 85.7 | 84.6 | 51.1 | 75.1 | — | — | — | |
| Diffusion PolicyHistory=false, 3D=false2026.06 | 78.3 | 92.5 | 68.3 | 50.5 | 72.4 | — | — | — | |
| Diffusion PolicyView Setting=Multi-view VLA, Category=General VLA2025.10 | 78.3 | 92.5 | 68.3 | 50.5 | 72.4 | — | — | — | |
| Open π0View Setting=Single-view VLA, Category=General VLA2025.10 | 77.2 | 84 | 83.6 | 66 | 77.7 | — | — | — | |
| UniACTModel Size=0.5B2025.06 | 77 | 87 | 77 | 70 | 77.8 | — | — | — | |
| SP-VLA (OpenVLA)Model Size=7B, Latency (ms)=124.262025.06 | 75.4 | 85.6 | 84.4 | 54.2 | 74.9 | — | — | — | |
| LAPAPretraining Method=Mix Videos2025.09 | 73.8 | 74.6 | 58.8 | 55.4 | 65.7 | — | — | — | |
| LatentLLMModel=X-VLA, Compression rate=0.12026.06 | 43 | 26 | 3.5 | 32.5 | 26.25 | — | — | 23.32 | |
| SVD-LLMModel=X-VLA, Compression rate=0.12026.06 | 41 | 26 | 4 | 31 | 25.5 | — | — | 22.6 | |
| ASVDModel=X-VLA, Compression rate=0.12026.06 | 32.5 | 21 | 5 | 13 | 17.88 | — | — | 15.38 | |
| Plain SVDModel=X-VLA, Compression rate=0.12026.06 | 7.5 | 1 | 0 | 2 | 2.63 | — | — | 1.72 |