Model-free learning on Stochastic MDPs Bilinear or BE
2Regret T ExponentDu et al. (2021), Jin et al. (2021), Xie et al. (2023)
Evaluation Results
| Method | Links | |
|---|---|---|
| Du et al. (2021), Jin et al. (2021), Xie et al. (2023)Exploration Mechanism=optimism, Estimation Protocol=on-policy2025.10 | 2 | |
| Du et al. (2021), Jin et al. (2021), Xie et al. (2023)Exploration Mechanism=optimism, Estimation Protocol=off-policy2025.10 | 2 | |
| Dig-DECExploration Mechanism=information gain, Estimation Protocol=on-policy2025.10 | 2 | |
| Foster et al. (2023b)Exploration Mechanism=information gain + optimism, Estimation Protocol=on-policy2025.10 | 3 | |
| Foster et al. (2023b)Exploration Mechanism=information gain + optimism, Estimation Protocol=off-policy2025.10 | 5 | |
| Dig-DECExploration Mechanism=information gain, Estimation Protocol=off-policy2025.10 | 7 |