Reinforcement Learning on Gymnasium HumanoidStandup
149,000Episodic ReturnCAPO
Evaluation Results
| Method | Links | |
|---|---|---|
| CAPOBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=true2026.03 | 149,000 | |
| Best-of-KBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=true2026.03 | 145,000 | |
| PPOBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=false2026.03 | 143,000 | |
| CAPO-AvgBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=true2026.03 | 141,000 | |
| TRPOBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=false2026.03 | 134,000 | |
| PPO-SWABackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=false2026.03 | 121,000 | |
| PPO-K×Backbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=true2026.03 | 104,000 |