Reinforcement Learning on D4RL Walker no right thigh (medium)
3,293Mean ReturnVGDF + BC
Evaluation Results
| Method | Links | |
|---|---|---|
| VGDF + BCtarget interaction steps=10^52023.05 | 3,293 | |
| H2Otarget interaction steps=10^52023.05 | 2,600 | |
| Offline onlybase algorithm=CQL, target interaction steps=10^52023.05 | 975 | |
| Symmetric samplingtarget interaction steps=10^52023.05 | 872 |