ResearchBenchmarksReinforcement Learning on Delayed MuJoCo Walker2d v2 (test)Follow5,274.89Max Average ReturnDE-MME3,999.03884,330.26944,661.54,992.7306Jun 19, 2021Evaluation ResultsMethodMethodLinksMax Average ReturnDE-MME2021.065,274.89MME2021.065,148.58MaxEnt(State)2021.064,430.61RND2021.064,098.63DAC2021.064,067.11SAC-Div2021.064,048.11