Safe Reinforcement Learning on Safety Gym POINTGOAL1 (original)
29Utility ScoreSEditor
Evaluation Results
| Method | Links | |
|---|---|---|
| SEditorSteps=2.5 x 10^72022.01 | 29 | |
| SEditorSteps=1 x 10^72022.01 | 27 | |
| Stooke et al. (2020)Steps=2.5 x 10^72022.01 | 26 | |
| SEditorSteps=5 x 10^62022.01 | 24 | |
| Stooke et al. (2020)Steps=1 x 10^72022.01 | 23 | |
| Stooke et al. (2020)Steps=5 x 10^62022.01 | 22 | |
| TRPO-LagSteps=1 x 10^72022.01 | 17 | |
| TRPO-LagSteps=5 x 10^62022.01 | 16 | |
| PPO-LagSteps=5 x 10^62022.01 | 14 | |
| PPO-LagSteps=1 x 10^72022.01 | 13 |