Reinforcement Learning on Procgen (train)
0.61Mean Normalized ScoreUCB-DrAC+PLR
Evaluation Results
| Method | Links | |
|---|---|---|
| UCB-DrAC+PLRTraining steps=200M2021.11 | 0.61 | |
| PLRTraining steps=200M2021.11 | 0.53 | |
| PPOTraining steps=200M2021.11 | 0.41 |
| Method | Links | |
|---|---|---|
| UCB-DrAC+PLRTraining steps=200M2021.11 | 0.61 | |
| PLRTraining steps=200M2021.11 | 0.53 | |
| PPOTraining steps=200M2021.11 | 0.41 |