Reinforcement Learning on D4RL HalfCheetah broken back thigh medium
5,761Mean ReturnH2O
Evaluation Results
| Method | Links | |
|---|---|---|
| H2Otarget interaction steps=10^52023.05 | 5,761 | |
| VGDF + BCtarget interaction steps=10^52023.05 | 4,834 | |
| Symmetric samplingtarget interaction steps=10^52023.05 | 2,439 | |
| Offline onlybase algorithm=CQL, target interaction steps=10^52023.05 | 1,128 |