ResearchBenchmarksReinforcement Learning on MiniGrid 8x8Follow93Clean RewardPPO91.9692.2392.592.77Jun 6, 2024Evaluation ResultsMethodMethodLinksClean RewardBest Attack RewardPPO2024.069312SA-PPOSmoothing type=smoothi...Smoothing type=smoothing, Time-discounting=false2024.069378TDRT-PPOSmoothing type=smoothi...Smoothing type=smoothing, Time-discounting=true2024.069274