Reinforcement Learning on Gymnasium Walker2d
3,871ReturnCAPO
Evaluation Results
| Method | Links | |
|---|---|---|
| CAPOBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=true2026.03 | 3,871 | |
| Best-of-KBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=true2026.03 | 3,770 | |
| TRPOBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=false2026.03 | 3,441 | |
| CAPO-AvgBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=true2026.03 | 3,223 | |
| PPOBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=false2026.03 | 2,518 | |
| PPO-K×Backbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=true2026.03 | 1,432 | |
| PPO-SWABackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=false2026.03 | 1,010 |