Reinforcement Learning on DMControl cartpole-swingup (random)
38.59Avg Normalized ScoreIQL+S2P
Evaluation Results
| Method | Links | |
|---|---|---|
| IQL+S2PBase Algorithm=IQL, S2P Augmentation=true2022.09 | 38.59 | |
| SLAC-offBase Algorithm=SLAC-off, S2P Augmentation=false2022.09 | 35.03 | |
| CQL+S2PBase Algorithm=CQL, S2P Augmentation=true2022.09 | 32.93 | |
| SLAC-off+S2PBase Algorithm=SLAC-off, S2P Augmentation=true2022.09 | 31.01 | |
| CQLBase Algorithm=CQL, S2P Augmentation=false2022.09 | 27.67 | |
| IQLBase Algorithm=IQL, S2P Augmentation=false2022.09 | 24.52 |