ResearchBenchmarksReinforcement Learning on MiniGrid 16x16Follow0.93Clean RewardTDRT-PPO0.91960.92230.9250.9277Jun 6, 2024Evaluation ResultsMethodMethodLinksClean RewardBest Attack RewardTDRT-PPOSmoothing type=smoothi...Smoothing type=smoothing, Time-discounting=true2024.060.930.74PPO2024.060.920.24SA-PPOSmoothing type=smoothi...Smoothing type=smoothing, Time-discounting=false2024.060.920.71