Reinforcement Learning on Humanoid v4
5,715RewardC-DSAC
Evaluation Results
| Method | Links | |
|---|---|---|
| C-DSACNumber of runs=100, Selection=Best models in training runs2026.04 | 5,715 | |
| PPO-ClipNumber of seeds=5, Evaluation window=last 10% of training2026.06 | 660 | |
| per-sample PPO-KLNumber of seeds=5, Evaluation window=last 10% of training2026.06 | 660 | |
| Adaptive βNumber of seeds=5, Evaluation window=last 10% of training2026.06 | 553 | |
| DDPG-AdaGammaBase RL Algorithm=DDPG, Discounting Strategy=AdaGamma2026.05 | 457.02 | |
| DDPG-UncertaintyBase RL Algorithm=DDPG, Discounting Strategy=Uncertainty2026.05 | 454.46 | |
| DDPG-CrossValidateBase RL Algorithm=DDPG, Discounting Strategy=CrossValidate2026.05 | 356.85 | |
| Fixed βNumber of seeds=5, Evaluation window=last 10% of training2026.06 | 342 | |
| TRPO-AdaGammaBase RL Algorithm=TRPO, Discounting Strategy=AdaGamma2026.05 | 284.49 | |
| TRPO-UncertaintyBase RL Algorithm=TRPO, Discounting Strategy=Uncertainty2026.05 | 250.65 | |
| TRPO-CrossValidateBase RL Algorithm=TRPO, Discounting Strategy=CrossValidate2026.05 | 221.1 | |
| TRPO-Fixed-γBase RL Algorithm=TRPO, Discounting Strategy=Fixed-γ2026.05 | 218.15 | |
| DDPG-Fixed-γBase RL Algorithm=DDPG, Discounting Strategy=Fixed-γ2026.05 | 165.82 |