Reinforcement Learning on Gymnasium HalfCheetah
2,737ReturnCAPO
Evaluation Results
| Method | Links | |
|---|---|---|
| CAPOBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=true2026.03 | 2,737 | |
| CAPO-AvgBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=true2026.03 | 1,967 | |
| Best-of-KBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=true2026.03 | 1,940 | |
| TRPOBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=false2026.03 | 1,629 | |
| PPOBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=false2026.03 | 1,604 | |
| PPO-K×Backbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=true2026.03 | 1,532 | |
| PPO-SWABackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=false2026.03 | 1,261 |