Reinforcement Learning on Gymnasium Hopper
3,455Total ReturnL2C2
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| L2C2RL Algorithm=SAC2026.01 | 3,455 | 0.893 | |
| ASAPRL Algorithm=SAC2026.01 | 3,448 | 0.498 | |
| CAPSRL Algorithm=SAC2026.01 | 3,413 | 0.793 | |
| SAC BaseRL Algorithm=SAC2026.01 | 3,349 | 0.856 | |
| CAPO-AvgBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=true2026.03 | 3,193 | — | |
| GRADRL Algorithm=SAC2026.01 | 3,190 | 0.588 | |
| PPO BaseBase RL Algorithm=PPO2026.01 | 2,902 | 1.709 | |
| GRADBase RL Algorithm=PPO2026.01 | 2,737 | 0.193 | |
| PPOBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=false2026.03 | 2,711 | — | |
| ASAPBase RL Algorithm=PPO2026.01 | 2,691 | 0.179 | |
| TRPOBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=false2026.03 | 2,583 | — | |
| CAPOBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=true2026.03 | 2,397 | — | |
| CAPSBase RL Algorithm=PPO2026.01 | 2,362 | 0.281 | |
| L2C2Base RL Algorithm=PPO2026.01 | 2,345 | 1.344 | |
| Best-of-KBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=true2026.03 | 2,255 | — | |
| PPO-SWABackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=false2026.03 | 2,240 | — | |
| PPO-K×Backbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=true2026.03 | 1,210 | — |