Continuous Control on MuJoCo Humanoid v2 (train)
6,242Mean ReturnTRGPPO
Evaluation Results
| Method | Links | |
|---|---|---|
| TRGPPOEnvironment steps=10M, Number of runs=52022.04 | 6,242 | |
| MCPOEnvironment steps=10M, Number of runs=52022.04 | 4,848 | |
| TRPOEnvironment steps=10M, Number of runs=52022.04 | 4,576 | |
| PPOEnvironment steps=10M, Number of runs=52022.04 | 3,375 | |
| MDPOEnvironment steps=10M, Number of runs=52022.04 | 1,620 | |
| Mean ψEnvironment steps=10M, Number of runs=5, N=402022.04 | 353 |