Reinforcement Learning on CartPole (Average Episode Reward)
499.11Average Episode RewardPPO-GBR
Evaluation Results
| Method | Links | |
|---|---|---|
| PPO-GBRrho (perturbation radius)=0, epsilon (uncertainty bound)=0.0032024.04 | 499.11 | |
| PPO-GBRrho (perturbation radius)=0, epsilon (uncertainty bound)=0.0052024.04 | 498.9 | |
| PPOrho (perturbation radius)=12024.04 | 498.53 | |
| PPOrho (perturbation radius)=02024.04 | 498.46 | |
| PPOrho (perturbation radius)=22024.04 | 498.35 | |
| PPO-GBRrho (perturbation radius)=0, epsilon (uncertainty bound)=0.0072024.04 | 497.64 | |
| PPO-PGDLCrho (perturbation radius)=3, epsilon (uncertainty bound)=0.0032024.04 | 494.07 | |
| PPO-GBRrho (perturbation radius)=0, epsilon (uncertainty bound)=0.0012024.04 | 485.16 | |
| PPOrho (perturbation radius)=32024.04 | 469.98 | |
| PPO-GBRrho (perturbation radius)=0, epsilon (uncertainty bound)=0.012024.04 | 417.98 | |
| PPO-PGDLCrho (perturbation radius)=4, epsilon (uncertainty bound)=0.0072024.04 | 417.69 | |
| PPO-PGDLCrho (perturbation radius)=5, epsilon (uncertainty bound)=0.0012024.04 | 357.11 | |
| PPOrho (perturbation radius)=42024.04 | 252.3 | |
| PPOrho (perturbation radius)=52024.04 | 223.29 |