Robot Manipulation on CALVIN (ABC->D)
4.75Average Successful LengthXiaomi-Robotics-0
Evaluation Results
| Method | Links | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Xiaomi-Robotics-0Setting=ABC→D2026.02 | 4.75 | — | — | — | — | — | 100 | 98.3 | 96 | 92.6 | 88.1 | |
| FLOWERSetting=ABC→D2026.02 | 4.53 | — | — | — | — | — | 99.4 | 95.8 | 90.7 | 84.9 | 77.8 | |
| DeFIView=Multi-View2026.03 | 4.51 | — | — | — | — | — | 97.9 | 94.2 | 90.7 | 87 | 81.2 | |
| UniVLASetting=ABC→D2026.02 | 4.41 | — | — | — | — | — | 98.9 | 94.8 | 89 | 82.8 | 75.1 | |
| BagelVLAKeyframe-forecasting=Enabled2026.02 | 4.405 | — | — | — | — | — | — | — | — | — | — | |
| VPPView=Multi-View2026.03 | 4.33 | — | — | — | — | — | 96.5 | 90.9 | 86.6 | 82 | 76.9 | |
| VPP2026.02 | 4.329 | — | — | — | — | — | — | — | — | — | — | |
| VPPSetting=ABC→D2026.02 | 4.29 | — | — | — | — | — | 95.7 | 91.2 | 86.3 | 81 | 75 | |
| Seer-LargeSetting=ABC→D2026.02 | 4.28 | — | — | — | — | — | 96.3 | 91.6 | 86.1 | 80.3 | 74 | |
| SeerView=Multi-View2026.03 | 4.28 | — | — | — | — | — | 96.3 | 91.6 | 86.1 | 80.3 | 74 | |
| RoboVLMsSetting=ABC→D2026.02 | 4.25 | — | — | — | — | — | 98 | 93.6 | 85.4 | 77.8 | 70.4 | |
| UP-VLAView=Multi-View2026.03 | 4.08 | — | — | — | — | — | 92.8 | 86.5 | 81.5 | 76.9 | 69.9 | |
| UP-VLA2026.02 | 4.078 | — | — | — | — | — | — | — | — | — | — | |
| DeFIView=Third View2026.03 | 4.05 | — | — | — | — | — | 92.9 | 87.2 | 81.2 | 75 | 68.4 | |
| GR-MGSetting=ABC→D2026.02 | 4.04 | — | — | — | — | — | 96.8 | 89.3 | 81.5 | 72.7 | 64.4 | |
| GR-MGInput=S-RGBD,G-RGBD,P2026.03 | 4.04 | — | — | — | — | — | 96.8 | 89.3 | 81.5 | 72.7 | 64.4 | |
| MoDESetting=ABC→D2026.02 | 4.01 | — | — | — | — | — | 96.2 | 88.9 | 81.1 | 71.8 | 63.5 | |
| GR00T N1*View=Multi-View2026.03 | 4.01 | — | — | — | — | — | 94.2 | 86.1 | 79.6 | 73.9 | 66.8 | |
| π0.5*View=Multi-View2026.03 | 3.97 | — | — | — | — | — | 94.8 | 87.4 | 78.2 | 71.7 | 64.3 | |
| π0*View=Multi-View2026.03 | 3.84 | — | — | — | — | — | 93.8 | 85 | 76.7 | 68.1 | 59.9 | |
| UniVLAView=Third View2026.03 | 3.8 | — | — | — | — | — | 95.5 | 85.8 | 75.4 | 66.9 | 56.5 | |
| Pretrained ProgressVLA (Full)Input=S-RGB2026.03 | 3.73 | — | — | — | — | — | 95.2 | 84.8 | 73.6 | 67.2 | 52 | |
| GHIL-GlueInput=S-RGB2026.03 | 3.69 | — | — | — | — | — | 95.2 | 88.5 | 73.2 | 62.5 | 49.8 | |
| Pretrained ProgressVLA (w/ cg(pretrained))Input=S-RGB2026.03 | 3.68 | — | — | — | — | — | 93.6 | 82 | 72 | 63.6 | 56.4 | |
| π02026.02 | 3.648 | — | — | — | — | — | — | — | — | — | — | |
| DitaInput=S-RGB2026.03 | 3.61 | — | — | — | — | — | 94.5 | 82.5 | 72.8 | 61.3 | 50 | |
| Pretrained ProgressVLA (w/ cg)Input=S-RGB2026.03 | 3.61 | — | — | — | — | — | 93.6 | 82.4 | 71.2 | 60.8 | 52.8 | |
| Pretrained ProgressVLA (w/o cg)Input=S-RGB2026.03 | 3.57 | — | — | — | — | — | 92.7 | 81.6 | 70.1 | 60.9 | 51.6 | |
| CLOVERView=Third View2026.03 | 3.53 | — | — | — | — | — | 96 | 83.5 | 70.8 | 57.5 | 45.4 | |
| VidmanView=Multi-View2026.03 | 3.42 | — | — | — | — | — | 91.5 | 76.4 | 68.2 | 59.2 | 46.7 | |
| 3DDASetting=ABC→D2026.02 | 3.35 | — | — | — | — | — | 93.8 | 80.3 | 66.2 | 53.3 | 41.2 | |
| BagelVLAKeyframe-forecasting=Disabled2026.02 | 3.345 | — | — | — | — | — | — | — | — | — | — | |
| 3D DiffuserInput=S-RGBD,G-RGBD,P,Cam2026.03 | 3.27 | — | — | — | — | — | 92.2 | 78.7 | 63.9 | 51.2 | 41.2 | |
| OpenVLAView=Multi-View2026.03 | 3.27 | — | — | — | — | — | 91.3 | 77.8 | 62 | 52.1 | 43.5 | |
| ProgressVLA (w/o cg)Input=S-RGB2026.03 | 3.24 | — | — | — | — | — | 89.4 | 76.8 | 63 | 52.2 | 43.1 | |
| UNILACTObservation input=Static RGB, Pre-training Dataset=OXE2026.02 | 3.1 | 0.897 | 0.729 | 0.584 | 0.493 | 0.401 | — | — | — | — | — | |
| GR-1Input=RGB+Proprio, Data=LANG, Foundation model=Video-pretrained Transformer2024.11 | 3.06 | — | — | — | — | — | — | — | — | — | — | |
| GR-1Observation input=(Static + Gripper) RGB + Proprio., Pre-training Dataset=CALVIN(ABC)2026.02 | 3.06 | 0.854 | 0.712 | 0.596 | 0.497 | 0.401 | — | — | — | — | — | |
| GR-1Setting=ABC→D2026.02 | 3.06 | — | — | — | — | — | 85.4 | 71.2 | 59.6 | 49.7 | 40.1 | |
| GR-1Input=S-RGB,G-RGB,P2026.03 | 3.06 | — | — | — | — | — | 85.4 | 71.2 | 59.6 | 49.7 | 40.1 | |
| GR-1View=Multi-View2026.03 | 3.06 | — | — | — | — | — | 85.4 | 71.2 | 59.6 | 49.7 | 40.1 | |
| DeeR w. onlineInput=RGB, Data=LANG, Foundation model=OpenFlamingo 3B, LLM GFLOPs=9.52024.11 | 2.9 | — | — | — | — | — | — | — | — | — | — | |
| UNILACTObservation input=Static RGB, Pre-training Dataset=CALVIN(ABC)2026.02 | 2.86 | 0.855 | 0.689 | 0.543 | 0.44 | 0.332 | — | — | — | — | — | |
| DeeRInput=RGB, Data=LANG, Foundation model=OpenFlamingo 3B, LLM GFLOPs=12.52024.11 | 2.82 | — | — | — | — | — | — | — | — | — | — | |
| SuSIEInput=RGB, Data=ALL, Foundation model=InstructPix2Pix [72]2024.11 | 2.69 | — | — | — | — | — | — | — | — | — | — | |
| SuSIEObservation input=Static RGB, Pre-training Dataset=CALVIN(ABC)2026.02 | 2.69 | 0.87 | 0.69 | 0.49 | 0.38 | 0.26 | — | — | — | — | — | |
| SuSIESetting=ABC→D2026.02 | 2.69 | — | — | — | — | — | 87 | 69 | 49 | 38 | 26 | |
| SuSIEInput=S-RGB2026.03 | 2.69 | — | — | — | — | — | 87 | 69 | 49 | 38 | 26 | |
| SuSIEView=Third View2026.03 | 2.69 | — | — | — | — | — | 87 | 69 | 49 | 38 | 26 | |
| MotoObservation input=Static RGB, Pre-training Dataset=CALVIN(ABC), Reproduced=true2026.02 | 2.6 | 0.827 | 0.637 | 0.485 | 0.374 | 0.278 | — | — | — | — | — | |
| RoboFlamingo++Input=RGB, Data=LANG, Foundation model=OpenFlamingo 3B, LLM GFLOPs=31.22024.11 | 2.59 | — | — | — | — | — | — | — | — | — | — | |
| Robo-FlamingoObservation input=(Static + Gripper) RGB, Pre-training Dataset=CALVIN(ABC)2026.02 | 2.48 | 0.824 | 0.619 | 0.466 | 0.331 | 0.235 | — | — | — | — | — | |
| RoboFlamingoSetting=ABC→D2026.02 | 2.48 | — | — | — | — | — | 82.4 | 61.9 | 46.6 | 33.1 | 23.5 | |
| RoboFlamingoInput=RGB, Data=LANG, Foundation model=OpenFlamingo 3B, LLM GFLOPs=31.22024.11 | 2.47 | — | — | — | — | — | — | — | — | — | — | |
| RoboFlamingoInput=S-RGB,G-RGB2026.03 | 2.47 | — | — | — | — | — | 82.4 | 61.9 | 46.6 | 33.1 | 23.5 | |
| MotoObservation input=Static RGB, Pre-training Dataset=OXE, Reproduced=true2026.02 | 2.4 | 0.824 | 0.6 | 0.439 | 0.316 | 0.228 | — | — | — | — | — | |
| SPILInput=RGB, Data=ALL, Foundation model=X2024.11 | 1.71 | — | — | — | — | — | — | — | — | — | — | |
| RT-1Input=RGB, Data=LANG, Foundation model=X2024.11 | 0.9 | — | — | — | — | — | — | — | — | — | — | |
| RT-1Observation input=Static RGB, Pre-training Dataset=CALVIN(ABC)2026.02 | 0.9 | 0.533 | 0.222 | 0.094 | 0.038 | 0.013 | — | — | — | — | — | |
| HULCInput=RGB, Data=ALL, Foundation model=X2024.11 | 0.67 | — | — | — | — | — | — | — | — | — | — | |
| HULCObservation input=Static RGB, Pre-training Dataset=CALVIN(ABC)2026.02 | 0.67 | 0.418 | 0.165 | 0.057 | 0.019 | 0.011 | — | — | — | — | — | |
| Diffusion PolicyObservation input=Static RGB, Pre-training Dataset=CALVIN(ABC)2026.02 | 0.56 | 0.402 | 0.123 | 0.026 | 0.008 | 0 | — | — | — | — | — |