Offline Reinforcement Learning on DMControl walker-walk (expert)
97.97Normalized ScoreCQL + S2P
Evaluation Results
| Method | Links | |
|---|---|---|
| CQL + S2PS2P Augmentation=true2022.09 | 97.97 | |
| CQLS2P Augmentation=false2022.09 | 95.43 | |
| IQL + S2PS2P Augmentation=true2022.09 | 94.97 | |
| IQLS2P Augmentation=false2022.09 | 94.34 | |
| LAWMaction-conditioned data percentage=5%2025.12 | 94.3 | |
| C-LAPaction-conditioned data percentage=100%, variant=Oracle2025.12 | 93.1 | |
| TD3+BCaction-conditioned data percentage=100%, variant=Oracle2025.12 | 92.8 | |
| IDM-TD3+BCaction-conditioned data percentage=5%2025.12 | 91 | |
| TD3+BCaction-conditioned data percentage=5%2025.12 | 89.6 | |
| C-LAPaction-conditioned data percentage=5%2025.12 | 88 | |
| SLAC-off + S2PS2P Augmentation=true2022.09 | 70.95 | |
| SLAC-offS2P Augmentation=false2022.09 | 11.71 |