Reinforcement Learning on HalfCheetah v4
10,554Max ReturnReLU
Evaluation Results
| Method | Links | |
|---|---|---|
| ReLUFrame=ANN2026.01 | 10,554 | |
| C-DSACNumber of runs=100, Selection=Best models in training runs2026.04 | 10,023 | |
| S-PLIFFrame=Vanilla2026.01 | 9,828 | |
| S-PLIFFrame=PT2026.01 | 9,405 | |
| PLIFFrame=Vanilla2026.01 | 9,252 | |
| PLIFFrame=PT2026.01 | 9,219 | |
| PDAEnvironment steps=1M, Number of seeds=10, Tests per epoch=102026.03 | 5,174.6 | |
| TRPOEnvironment steps=1M, Number of seeds=10, Tests per epoch=102026.03 | 4,496.8 | |
| PPOEnvironment steps=1M, Number of seeds=10, Tests per epoch=102026.03 | 4,067.5 | |
| NPGEnvironment steps=1M, Number of seeds=10, Tests per epoch=102026.03 | 3,556.4 | |
| Adaptive βNumber of seeds=5, Evaluation window=last 10% of training2026.06 | 2,207 | |
| Fixed βNumber of seeds=5, Evaluation window=last 10% of training2026.06 | 2,111 | |
| PPO-ClipNumber of seeds=5, Evaluation window=last 10% of training2026.06 | 1,955 | |
| per-sample PPO-KLNumber of seeds=5, Evaluation window=last 10% of training2026.06 | 1,955 |