Robust Reinforcement Learning on Average-Reward Markov Decision Processes
2Sample ComplexityChen et al.
Evaluation Results
| Method | Links | |
|---|---|---|
| Chen et al.Learning Type=Model-Based, AMDP Structure=Uniformly ergodic, Uncertainty Set=KL2025.05 | 2 | |
| Roch et al.Learning Type=Model-Based, AMDP Structure=Unichain, Uncertainty Set=TV2025.05 | 2 | |
| Xu et al. (2025b)Learning Type=Model-Free, AMDP Structure=Irreducible & aperiodic, Uncertainty Set=TV, Evaluation Protocol=policy evaluation2025.05 | 2 | |
| Xu et al. (2025a)Learning Type=Model-Free, AMDP Structure=Irreducible & aperiodic, Uncertainty Set=TV2025.05 | 2 | |
| Robust Halpern Iteration (RHI)Learning Type=Model-Free, AMDP Structure=Irreducible, Uncertainty Set=KL/CS2025.05 | 2 | |
| Robust Halpern Iteration (RHI)Learning Type=Model-Free, AMDP Structure=Irreducible, Uncertainty Set=Contamination2025.05 | 2 |