ResearchTasksOptimal Policy EstimationFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedContinuous Simulation Setting epsilon = 0.9Super0.06Mean Regret3Jun 4, 2026Continuous Simulation Setting epsilon = 0.7Super0.1Mean Regret3Jun 4, 2026Continuous Simulation Setting (epsilon = 0.5)SZonly0.11Mean Regret3Jun 4, 2026