ResearchBenchmarksSafe Reinforcement Learning on single-gate Double Integrator dynamicsFollow0Mean Safety Violations per EpisodeTD3-0.42.357.7May 22, 2024Evaluation ResultsMethodMethodLinksMean Safety Violations per EpisodeMean Shield Invocations per EpisodeStd Dev Shield Invocations per EpisodeStd Dev Safety Violations per EpisodeTD32024.050——0CPO2024.052——1.2PPO-lag2024.0510——5.5DMPS2024.05—0.10—MPS2024.05—0.20.1—