Offline Reinforcement Learning on DMControl cheetah-run (expert)
96.28Normalized ScoreCQL + S2P
Evaluation Results
| Method | Links | |
|---|---|---|
| CQL + S2PS2P Augmentation=true2022.09 | 96.28 | |
| CQLS2P Augmentation=false2022.09 | 94.2 | |
| IQL + S2PS2P Augmentation=true2022.09 | 87.18 | |
| IDM-TD3+BCaction-conditioned data percentage=5%2025.12 | 85.6 | |
| TD3+BCaction-conditioned data percentage=100%, variant=Oracle2025.12 | 83.4 | |
| IQLS2P Augmentation=false2022.09 | 79.89 | |
| C-LAPaction-conditioned data percentage=100%, variant=Oracle2025.12 | 79.6 | |
| C-LAPaction-conditioned data percentage=5%2025.12 | 69.1 | |
| TD3+BCaction-conditioned data percentage=5%2025.12 | 52.9 | |
| LAWMaction-conditioned data percentage=5%2025.12 | 52.4 | |
| SLAC-off + S2PS2P Augmentation=true2022.09 | 14.41 | |
| SLAC-offS2P Augmentation=false2022.09 | 8.92 |