ResearchBenchmarksSafe Reinforcement Learning on dynamic-obst Differential Drive dynamicsFollow0.8Mean Safety Violations per EpisodeTD30.7161.2831.852.417May 22, 2024Evaluation ResultsMethodMethodLinksMean Safety Violations per EpisodeMean Shield Invocations per EpisodeSD Shield Invocations per EpisodeSD Safety Violations per EpisodeTD32024.050.8——0.16CPO2024.051.7——1.9PPO-lag2024.052.9——2.2DMPS2024.05—105.239.9—MPS2024.05—144.939.9—