Finding an ε-optimal policy from a single trajectory on Weakly communicating average-reward MDPs
-2Sample Complexity (Exponent on ε)Our work
Evaluation Results
| Method | Links | |
|---|---|---|
| Our workMethod type=model-free, Additional assumptions=none2026.06 | -2 | |
| Lee et al.Method type=model-free, Additional assumptions=finite coverage & ergodicity on behavior policy2026.06 | -8 |