Offline-to-online Reinforcement Learning on Robomimic multi-human (MH)
100Lift Success RateQC-FQL
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| QC-FQLinference_type=distilled one-step, action_chunking=true2026.05 | 100 | 72 | 94.4 | |
| DFPinference_type=teacher-free one-step, action_chunking=true2026.05 | 100 | 93.2 | 90.6 | |
| MVPinference_type=teacher-free one-step, action_chunking=true2026.05 | 99.8 | 79.4 | 83.6 | |
| QC-BFNinference_type=multi-step, action_chunking=true2026.05 | 99.6 | 88.4 | 90.6 | |
| BFNinference_type=multi-step, action_chunking=false2026.05 | 97.6 | 32.8 | 82 | |
| FQLinference_type=distilled one-step, action_chunking=false2026.05 | 96.8 | 10.8 | 58.4 |