Reinforcement Learning on Gymnasium Humanoid
6,367ReturnCAPO
Evaluation Results
| Method | Links | |
|---|---|---|
| CAPOBackbone=(64, 64) MLP, Training steps=4M, Seeds=8, Compute-matched=true2026.03 | 6,367 | |
| TRPOBackbone=(64, 64) MLP, Training steps=4M, Seeds=8, Compute-matched=false2026.03 | 3,730 | |
| PPO-SWABackbone=(64, 64) MLP, Training steps=4M, Seeds=8, Compute-matched=false2026.03 | 750 | |
| PPOBackbone=(64, 64) MLP, Training steps=4M, Seeds=8, Compute-matched=false2026.03 | 739 | |
| Best-of-KBackbone=(64, 64) MLP, Training steps=4M, Seeds=8, Compute-matched=true2026.03 | 716 | |
| CAPO-AvgBackbone=(64, 64) MLP, Training steps=4M, Seeds=8, Compute-matched=true2026.03 | 696 | |
| PPO-K×Backbone=(64, 64) MLP, Training steps=4M, Seeds=8, Compute-matched=true2026.03 | 429 |