Language-conditioned Robotic Instruction Following on CALVIN ABC→D
98.9Success Rate (1 Task)Unified-VLA
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Unified-VLAIntermediate=future visual states, Pretrain=true2026.05 | 98.9 | 94.8 | 89 | 82.8 | 75.1 | 4.41 | |
| DreamVLAIntermediate=future visual states, Pretrain=true2026.05 | 98.2 | 94.6 | 89.5 | 83.4 | 78.1 | 4.44 | |
| OASISIntermediate=SE(3)-supervised features, Pretrain=false2026.05 | 98.1 | 94.9 | 91.7 | 88.9 | 83.3 | 4.57 | |
| VPPIntermediate=future visual states, Pretrain=true2026.05 | 96.5 | 90.9 | 86.6 | 82 | 76.9 | 4.33 | |
| Seer-LargeIntermediate=future visual states, Pretrain=true2026.05 | 96.3 | 91.6 | 86.1 | 80.3 | 74 | 4.28 | |
| ReconVLAIntermediate=spatial features, Pretrain=true2026.05 | 95.6 | 87.6 | 76.9 | 69.3 | 64.1 | 3.95 | |
| 3D Diffuser ActorIntermediate=3D feature, Pretrain=false, Input Modality=multi-view RGB-D2026.05 | 93.8 | 80.3 | 66.2 | 53.3 | 41.2 | 3.35 | |
| SuSIEIntermediate=future visual states, Pretrain=true2026.05 | 87 | 69 | 49 | 38 | 26 | 2.69 |