Reinforcement Learning on DMControl walker-walk (expert)
97.97Averaged Normalized ScoreCQL+S2P
Evaluation Results
| Method | Links | |
|---|---|---|
| CQL+S2PBase Algorithm=CQL, S2P Augmentation=true2022.09 | 97.97 | |
| CQLBase Algorithm=CQL, S2P Augmentation=false2022.09 | 95.43 | |
| IQL+S2PBase Algorithm=IQL, S2P Augmentation=true2022.09 | 94.97 | |
| IQLBase Algorithm=IQL, S2P Augmentation=false2022.09 | 94.34 | |
| SLAC-off+S2PBase Algorithm=SLAC-off, S2P Augmentation=true2022.09 | 70.95 | |
| SLAC-offBase Algorithm=SLAC-off, S2P Augmentation=false2022.09 | 11.71 |