ResearchBenchmarksReinforcement Learning on DI-double-gatesFollow12.7Mean ScoreDMPS-7.476-2.23838.238May 22, 2024Evaluation ResultsMethodMethodLinksMean ScoreSDDMPS2024.0512.71MPS2024.0511.51.1TD32024.05-3.61PPO-Lag2024.05-3.911.1CPO2024.05-6.75.4