Reinforcement Learning on DMControl ball-in-cup-catch (mixed)
51.28Averaged Normalized ScoreCQL+S2P
Evaluation Results
| Method | Links | |
|---|---|---|
| CQL+S2PBase Algorithm=CQL, S2P Augmentation=true2022.09 | 51.28 | |
| IQLBase Algorithm=IQL, S2P Augmentation=false2022.09 | 41.94 | |
| SLAC-off+S2PBase Algorithm=SLAC-off, S2P Augmentation=true2022.09 | 40.41 | |
| IQL+S2PBase Algorithm=IQL, S2P Augmentation=true2022.09 | 37.79 | |
| CQLBase Algorithm=CQL, S2P Augmentation=false2022.09 | 30.82 | |
| SLAC-offBase Algorithm=SLAC-off, S2P Augmentation=false2022.09 | 28.54 |