Reinforcement Learning on Continuous Cartpole
542.38Average Final RewardBaseline PPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Baseline PPOUsable?=No: Constraints2025.10 | 542.38 | 77.3 | |
| Reversibility Aware ControlUsable?=No: Constraints2025.10 | 539.15 | 35.1 | |
| PPO with Safety Oracle shieldUsable?=No: Oracle2025.10 | 537.76 | 0 | |
| PPO with SAVMPC shieldUsable?=Yes2025.10 | 533.66 | 0 | |
| Leave No TraceUsable?=No: Constraints2025.10 | 119.56 | 58.5 |