Reinforcement Learning on InvertedPendulum v4
1,000Average Episodic RewardNPG
Evaluation Results
| Method | Links | |
|---|---|---|
| NPGEnvironment steps=1M, Number of seeds=10, Tests per epoch=102026.03 | 1,000 | |
| PDAEnvironment steps=1M, Number of seeds=10, Tests per epoch=102026.03 | 1,000 | |
| PPOEnvironment steps=1M, Number of seeds=10, Tests per epoch=102026.03 | 993.3 | |
| TRPOEnvironment steps=1M, Number of seeds=10, Tests per epoch=102026.03 | 380.3 |