Robotic Manipulation on SIMPLER-Bridge
56.3Average Success Ratemimic-video
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| mimic-videoTraining regime=scratch, Inference-time policy optimization=per task τv-tuning, Inputs=third-person image, language instruction, robot proprioceptive state (optional)2025.12 | 56.3 | 54.2 | 41.7 | 29.2 | 100 | |
| OpenVLA2026.03 | 53.7 | — | — | — | — | |
| mimic-videoTraining regime=scratch, Inference-time policy optimization=none, Inputs=third-person image, language instruction, robot proprioceptive state (optional)2025.12 | 46.9 | 37.5 | 37.5 | 12.5 | 100 | |
| OmniStream2026.03 | 45.8 | — | — | — | — | |
| FLOWERTraining regime=finetuned, Inputs=third-person image, language instruction, robot proprioceptive state (optional)2025.12 | 45 | 13 | 71 | 8 | 88 | |
| ThinkActTraining regime=pretrained, Inputs=third-person image, language instruction, robot proprioceptive state (optional)2025.12 | 43.8 | 37.5 | 58.3 | 8.7 | 70.8 | |
| π0.5-style VLATraining regime=scratch, Inputs=third-person image, language instruction, robot proprioceptive state (optional)2025.12 | 35.4 | 25 | 29.2 | 20.8 | 66.7 | |
| OctoTraining regime=finetuned, Inputs=third-person image, language instruction, robot proprioceptive state (optional)2025.12 | 16 | 8.3 | 12.5 | 0 | 43.1 | |
| OpenVLATraining regime=finetuned, Inputs=third-person image, language instruction, robot proprioceptive state (optional)2025.12 | 14.6 | 4.2 | 8.3 | 0 | 45.8 |