ResearchTasksInfinite-horizon average-reward Reinforcement LearningFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedInfinite-horizon Average-reward Unichain CMDPsAgarwal et al.0Constraint Violation2Feb 26, 2026