Robot Manipulation on SimplerEnv WidowX Robot tasks (test)
89.6Success Rate (Spoon)LangForce
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| LangForcebackbone=Qwen3-VL-4B2026.01 | 89.6 | — | — | 63.8 | — | 33.3 | — | 79.2 | 66.5 | — | |
| QwenGR00Tmode=Baseline, backbone=Qwen3-VL-4B2026.01 | 87.5 | — | — | 50 | — | 29.2 | — | 54.2 | 55.2 | — | |
| NORA-Long RFT (Ours)Action Model Type=Discrete, Training Strategy=RFT2026.02 | 84.3 | — | — | 50.2 | — | 64.4 | — | 77 | — | 69 | |
| NORA-Long (Baseline)Action Model Type=Discrete, Training Strategy=SFT2026.02 | 80.2 | — | — | 46 | — | 60.3 | — | 75.7 | — | 65.5 | |
| ST4VLACo-Train=true2026.02 | 80.2 | — | — | 79.2 | — | 35.4 | — | 98 | — | 73.2 | |
| Dream-VLA2025.12 | 79.2 | 91.7 | 58.3 | 41.7 | 79.2 | 20.8 | 100 | 100 | 71.4 | 60.4 | |
| GR00T N1.5Co-Train=false, Category=Visual Matching2026.02 | 75.3 | — | — | 54.3 | — | 57 | — | 61.3 | — | 61.9 | |
| VideoVLA2026.01 | 75 | — | — | 20.8 | — | 45.8 | — | 70.8 | 53.1 | — | |
| VLA-JEPAtraining_data=subsets of the OXE dataset, setting=visual matching2026.02 | 75 | — | — | 70.8 | — | 12.5 | — | 70.8 | — | 57.3 | |
| VLA-JEPA (w/o human videos)training_data=subsets of the OXE dataset, setting=visual matching, ablation=w/o human videos2026.02 | 75 | — | — | 54.2 | — | 20.8 | — | 79.2 | — | 57.3 | |
| CogACT2026.01 | 71.7 | — | — | 50.8 | — | 15 | — | 67.5 | 51.3 | — | |
| CogACTCo-Train=false, Category=Visual Matching2026.02 | 71.7 | — | — | 50.8 | — | 15 | — | 67.5 | — | 51.3 | |
| CogACTRobot type=WidowX, Evaluation protocol=Visual Matching2026.01 | 71.7 | — | — | 50.8 | — | 15 | — | 67.5 | — | 51.3 | |
| AC2-VLARobot type=WidowX, Evaluation protocol=Visual Matching2026.01 | 71.2 | — | — | 58 | — | 14.8 | — | 74 | — | 54.5 | |
| LAPAtraining_data=in-distribution expert demonstrations collected in simulation environment, setting=visual matching2026.02 | 70.8 | — | — | 45.8 | — | 54.2 | — | 58.3 | — | 57.3 | |
| Vanilla Co-training VLACo-Train=true2026.02 | 70.3 | — | — | 68.4 | — | 20.5 | — | 85.2 | — | 61.1 | |
| Isaac-GR00T-N1.6-Bridge2026.01 | 64.5 | — | — | 65.5 | — | 5.5 | — | 93 | 57.1 | — | |
| π0Action Model Type=Continuous, Training Strategy=SFT2026.02 | 63.3 | — | — | 58.8 | — | 21.3 | — | 79.2 | — | 55.7 | |
| GR00T-N12025.12 | 62.5 | 83.3 | 54.2 | 45.8 | 70.8 | 16.7 | 41.7 | 20.8 | 49.5 | 36.5 | |
| ThinkActAction Model Type=Continuous, Training Strategy=SFT + RFT2026.02 | 58.3 | — | — | 37.5 | — | 8.7 | — | 70.8 | — | 43.8 | |
| LLaDA-VLA2025.12 | 56.9 | — | — | 76.3 | — | 30.6 | — | 58.3 | — | 55.5 | |
| Vanilla VLACo-Train=false2026.02 | 56.6 | — | — | 63.3 | — | 27 | — | 71.8 | — | 54.7 | |
| pi_0 + TACOframework=TACO, type=test-time scaling2025.12 | 52 | — | — | 52 | — | 30 | — | 88 | — | 55.5 | |
| RoboVLM2026.01 | 50 | — | — | 37.5 | — | 0 | — | 83.3 | 42.7 | — | |
| π0.52026.01 | 49.3 | — | — | 64.7 | — | 44.7 | — | 69.7 | 57.1 | — | |
| villa-xtraining_data=subsets of the OXE dataset, setting=visual matching2026.02 | 48.3 | — | — | 24.2 | — | 19.2 | — | 71.7 | — | 40.8 | |
| RoboVLMstraining_data=subsets of the OXE dataset, setting=visual matching2026.02 | 45.8 | — | — | 20.8 | — | 4.2 | — | 79.2 | — | 37.5 | |
| Octo-SmallModel size=Small2026.01 | 41.7 | — | — | 8.2 | — | 0 | — | 56.7 | 26.7 | — | |
| Octo-SmallCo-Train=false, Category=Visual Matching2026.02 | 41.7 | — | — | 8.2 | — | 0 | — | 56.7 | — | 26.7 | |
| Octo-SmallRobot type=WidowX, Evaluation protocol=Visual Matching2026.01 | 41.7 | — | — | 8.2 | — | 0 | — | 56.7 | — | 26.7 | |
| Magma2026.01 | 37.5 | — | — | 29.2 | — | 20.8 | — | 91.7 | 44.8 | — | |
| MagmaCo-Train=true, Category=Visual Matching2026.02 | 37.5 | — | — | 31 | — | 12.7 | — | 60.5 | — | 35.8 | |
| pi_02025.12 | 36 | — | — | 42 | — | 34 | — | 80 | — | 48 | |
| OpenVLA-OFT2026.01 | 34.2 | — | — | 30 | — | 30 | — | 72.5 | 41.8 | — | |
| OpenVLA-OFTtraining_data=subsets of the OXE dataset, setting=visual matching2026.02 | 34.2 | — | — | 30 | — | 30 | — | 72.5 | — | 41.8 | |
| Octo-Base2025.12 | 33 | 50 | 50 | 25 | 29.2 | 0 | 40 | 23.3 | 31.3 | 20.3 | |
| π02026.01 | 29.2 | — | — | 62.5 | — | 29.2 | — | 91.6 | 53.1 | — | |
| DiscreteDiffusionVLA2025.12 | 29.2 | 70.8 | 58.3 | 29.2 | 62.5 | 20.8 | 91.7 | 70.8 | 54.2 | 37.5 | |
| RoboVLMAction Model Type=Continuous, Training Strategy=SFT2026.02 | 29.2 | — | — | 25 | — | 12.5 | — | 58.3 | — | 31.3 | |
| RoboVLM2025.12 | 29.2 | — | — | 25 | — | 12.5 | — | 58.3 | — | 31.3 | |
| π02025.12 | 29.1 | 45.8 | 25 | 0 | 50 | 16.6 | 91.6 | 62.5 | 40.1 | 27.1 | |
| π0 + FASTSpeed Config=FAST2025.12 | 29.1 | 62.5 | 58.5 | 21.9 | 54 | 10.8 | 83.3 | 66.6 | 48.3 | 32.1 | |
| π0Co-Train=false, Category=Visual Matching2026.02 | 29.1 | — | — | 0 | — | 16.6 | — | 62.5 | — | 27.1 | |
| π0-FASTCo-Train=false, Category=Visual Matching2026.02 | 29.1 | — | — | 21.9 | — | 10.8 | — | 66.6 | — | 48.3 | |
| π0training_data=subsets of the OXE dataset, setting=visual matching2026.02 | 29.1 | — | — | 0 | — | 16.6 | — | 62.5 | — | 40.1 | |
| π0-Fasttraining_data=subsets of the OXE dataset, setting=visual matching2026.02 | 29.1 | — | — | 21.9 | — | 10.8 | — | 66.7 | — | 48.3 | |
| π0-FASTAction Model Type=Discrete, Training Strategy=SFT2026.02 | 29 | — | — | 22 | — | 83 | — | 48 | — | 45.5 | |
| SpatialVLA2026.01 | 20.8 | — | — | 20.8 | — | 25 | — | 70.8 | 34.4 | — | |
| RoboVLM2025.12 | 20.8 | 37.5 | 33.3 | 25 | 8.3 | 8.3 | 0 | 0 | 16.7 | 13.5 | |
| SpatialVLA2025.12 | 16.7 | 20.8 | 29.2 | 25 | 62.5 | 29.2 | 100 | 100 | 47.9 | 42.7 | |
| SpatialVLAAction Model Type=Discrete, Training Strategy=SFT2026.02 | 16.7 | — | — | 25 | — | 29.2 | — | 100 | — | 42.7 | |
| SpatialVLACo-Train=false, Category=Visual Matching2026.02 | 16.7 | — | — | 25 | — | 29.2 | — | 100 | — | 42.7 | |
| SpatialVLA2025.12 | 16.7 | — | — | 25 | — | 29.2 | — | 100 | — | 42.7 | |
| Octo-BaseModel size=Base2026.01 | 15.8 | — | — | 12.5 | — | 0 | — | 41.7 | 17.5 | — | |
| Octo-BaseCo-Train=false, Category=Visual Matching2026.02 | 15.8 | — | — | 12.5 | — | 0 | — | 41.7 | — | 17.5 | |
| Octo-BaseRobot type=WidowX, Evaluation protocol=Visual Matching2026.01 | 15.8 | — | — | 12.5 | — | 0 | — | 41.7 | — | 17.5 | |
| TraceVLA2026.01 | 12.5 | — | — | 16.6 | — | 16.6 | — | 65 | 27.7 | — | |
| OpenVLA-OFTFine-tuned=OFT2025.12 | 12.5 | 50 | 41.7 | 4.2 | 70.8 | 20.8 | 91.7 | 37.5 | 41.2 | 18.8 | |
| Octo-BaseAction Model Type=Continuous, Training Strategy=SFT2026.02 | 12.5 | — | — | 8.3 | — | 0 | — | 43.1 | — | 16 | |
| Octo2025.12 | 12.5 | — | — | 8.3 | — | 0 | — | 43.1 | — | 16 | |
| OpenVLACo-Train=false, Category=Visual Matching2026.02 | 4.2 | — | — | 0 | — | 0 | — | 12.5 | — | 4.2 | |
| OpenVLARobot type=WidowX, Evaluation protocol=Visual Matching2026.01 | 4.2 | — | — | 0 | — | 0 | — | 12.5 | — | 4.2 | |
| GR00T N1training_data=subsets of the OXE dataset, setting=visual matching2026.02 | 1.4 | — | — | 0 | — | 0 | — | 13.9 | — | 3.8 | |
| Octo-Smallevaluation_protocol=zero-shot, pre-training_dataset=OXE dataset [13], model_size=Small2025.01 | 0.472 | 0.778 | 0.278 | 0.097 | 0.403 | 0.042 | 87.5 | 56.9 | 0.3 | — | |
| RoboVLMevaluation_protocol=fine-tuning, fine-tuning_dataset=BridgeData V2 [64]2025.01 | 0.292 | 0.542 | 0.25 | 0.25 | 0.458 | 0.125 | 58.3 | 58.3 | 0.313 | — | |
| RoboVLMevaluation_protocol=zero-shot, pre-training_dataset=OXE dataset [13]2025.01 | 0.208 | 0.375 | 0.333 | 0.25 | 0.083 | 0.083 | 0 | 0 | 0.135 | — | |
| SpatialVLAevaluation_protocol=zero-shot, pre-training_dataset=OXE dataset [13]2025.01 | 0.208 | 0.25 | 0.417 | 0.208 | 0.583 | 0.25 | 79.2 | 70.8 | 0.344 | — | |
| SpatialVLAevaluation_protocol=fine-tuning, fine-tuning_dataset=BridgeData V2 [64]2025.01 | 0.167 | 0.208 | 0.292 | 0.25 | 0.625 | 0.292 | 100 | 100 | 0.427 | — | |
| Octo-Baseevaluation_protocol=zero-shot, pre-training_dataset=OXE dataset [13], model_size=Base2025.01 | 0.125 | 0.347 | 0.528 | 0.083 | 0.319 | 0 | 66.7 | 43.1 | 0.16 | — | |
| RT-1-Xevaluation_protocol=zero-shot, pre-training_dataset=OXE dataset [13]2025.01 | 0 | 0.167 | 0.208 | 0.042 | 0.083 | 0 | 0 | 0 | 0.011 | — | |
| OpenVLAevaluation_protocol=zero-shot, pre-training_dataset=OXE dataset [13]2025.01 | 0 | 0.041 | 0.333 | 0 | 0.125 | 0 | 8.3 | 4.1 | 0.01 | — | |
| RT-1-X2026.01 | 0 | — | — | 4.2 | — | 0 | — | 0 | 1.1 | — | |
| RT-1-X2025.12 | 0 | 4.2 | 16.7 | 0 | 0 | 0 | 3.3 | 0 | 3 | 0 | |
| OpenVLA2025.12 | 0 | 4.1 | 33 | 0 | 12.5 | 0 | 8.3 | 4.1 | 7.8 | 1 | |
| RT-1-XAction Model Type=Discrete, Training Strategy=SFT2026.02 | 0 | — | — | 4.2 | — | 0 | — | 0 | — | 1.1 | |
| OpenVLAAction Model Type=Discrete, Training Strategy=SFT2026.02 | 0 | — | — | 0 | — | 0 | — | 4.1 | — | 1 | |
| RT-1-XCo-Train=false, Category=Visual Matching2026.02 | 0 | — | — | 4.2 | — | 0 | — | 0 | — | 1.1 | |
| RT-1-XRobot type=WidowX, Evaluation protocol=Visual Matching2026.01 | 0 | — | — | 4.2 | — | 0 | — | 0 | — | 1.1 | |
| RT-1-X2025.12 | 0 | — | — | 4.2 | — | 0 | — | 0 | — | 1.1 | |
| UniVLAtraining_data=subsets of the OXE dataset, setting=visual matching2026.02 | — | — | — | — | — | — | — | — | — | 42.7 |