Reinforcement Learning on Gymnasium Ant
4,966Cumulative ReturnASAP
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ASAPRL Algorithm=SAC2026.01 | 4,966 | 1.226 | |
| CAPSRL Algorithm=SAC2026.01 | 4,532 | 1.739 | |
| SAC BaseRL Algorithm=SAC2026.01 | 4,239 | 1.715 | |
| L2C2RL Algorithm=SAC2026.01 | 4,125 | 1.811 | |
| GRADRL Algorithm=SAC2026.01 | 3,938 | 1.423 | |
| ASAPBase RL Algorithm=PPO2026.01 | 2,574 | 1.092 | |
| GRADBase RL Algorithm=PPO2026.01 | 2,419 | 1.121 | |
| CAPOBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=true2026.03 | 2,328 | — | |
| CAPSBase RL Algorithm=PPO2026.01 | 2,185 | 1.259 | |
| PPOBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=false2026.03 | 2,142 | — | |
| CAPO-AvgBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=true2026.03 | 1,738 | — | |
| Best-of-KBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=true2026.03 | 1,664 | — | |
| PPO BaseBase RL Algorithm=PPO2026.01 | 1,434 | 1.497 | |
| L2C2Base RL Algorithm=PPO2026.01 | 1,393 | 1.459 | |
| TRPOBackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=false2026.03 | 1,029 | — | |
| PPO-SWABackbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=false2026.03 | 943 | — | |
| PPO-K×Backbone=(64, 64) MLP, Training steps=1M, Seeds=8, Compute-matched=true2026.03 | 232 | — |