ResearchTasksEpsilon-optimal Policy IdentificationFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedAverage-reward Markov Decision Process (AMDP)——Primary metric0Feb 26, 2026