ResearchBenchmarksReinforcement Learning on Grid World Npick=3 Dense (test)Follow3Max Average ReturnLOVE1.962.232.52.77Dec 8, 2022Evaluation ResultsMethodMethodLinksMax Average ReturnLOVE2022.123Option-Criticnumber of options=8number of options=82022.122