Reinforcement Learning on Humanoid 3M v4
5,020Average Episodic RewardPDA
Evaluation Results
| Method | Links | |
|---|---|---|
| PDAEnvironment steps=3M, Number of seeds=10, Tests per epoch=102026.03 | 5,020 | |
| TRPOEnvironment steps=3M, Number of seeds=10, Tests per epoch=102026.03 | 4,745.1 | |
| NPGEnvironment steps=3M, Number of seeds=10, Tests per epoch=102026.03 | 4,650.6 | |
| PPOEnvironment steps=3M, Number of seeds=10, Tests per epoch=102026.03 | 933.3 |