ResearchBenchmarksReinforcement Learning on MuJoCo HalfCheetahFollow3,680,220Clean RewardPPO2,934,725.123,128,267.063,321,8093,515,350.94Jun 6, 2024Evaluation ResultsMethodMethodLinksClean RewardBest Attack RewardPPOSmoothing=false, Time-...Smoothing=false, Time-discounting=false2024.063,680,220160,540TDRT-PPOSmoothing=true, Time-d...Smoothing=true, Time-discounting=true2024.063,577,1972,981,462SA-PPOSmoothing=true, Time-d...Smoothing=true, Time-discounting=false2024.062,963,3982,826,513