Reinforcement Learning on Ant (e1, e2, Av. Reward)
7,010.88Average RewardPPO-PGDLC
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| PPO-PGDLCrho (perturbation radius)=0, epsilon (uncertainty bound)=0.0012024.04 | 7,010.88 | — | — | |
| PPO-PGDLCrho (perturbation radius)=1, epsilon (uncertainty bound)=0.0032024.04 | 6,885.73 | — | — | |
| PPOrho (perturbation radius)=02024.04 | 6,339.91 | — | — | |
| PPO-PGDLCrho (perturbation radius)=2, epsilon (uncertainty bound)=0.0032024.04 | 6,262.51 | — | — | |
| PPO-GBRrho (perturbation radius)=0, epsilon (uncertainty bound)=0.0012024.04 | 5,553.56 | — | — | |
| PPO-PGDLCrho (perturbation radius)=3, epsilon (uncertainty bound)=0.0032024.04 | 5,539.27 | — | — | |
| PPO-PGDLCrho (perturbation radius)=4, epsilon (uncertainty bound)=0.0032024.04 | 4,794.98 | — | — | |
| PPO-PGDLCrho (perturbation radius)=5, epsilon (uncertainty bound)=0.0032024.04 | 4,223.23 | — | — | |
| Curriculum Adversarial TrainingPhase=22026.06 | 0.96 | 0.09 | 0.24 | |
| Scenario Generation for Risk-Aware Reinforcement LearningPhase=22026.06 | 0.94 | 0.09 | 0.18 | |
| CoDEPhase=22026.06 | 0.93 | 0.08 | 0.29 | |
| Genetic Curriculum LearningPhase=22026.06 | 0.9 | 0.14 | 0.45 | |
| Vanilla PPOPhase=22026.06 | 0.81 | 0.2 | 0.45 | |
| Epsilon Greedy PPOPhase=22026.06 | 0.77 | 0.22 | 0.48 | |
| Perturbed ExplorationPhase=22026.06 | 0.74 | 0.12 | 0.51 | |
| Scenario Generation for Risk-Aware Reinforcement LearningPhase=12026.06 | 0.62 | 0.11 | 0.41 |