Regret Bound Estimation on Hybrid MDPs (Bilinear or Coverable, Bellman Complete, On-Policy)
3Regret Bound (On-Policy)Dig-DEC
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Dig-DECModel-Free=✓, Bandit Feedback=✓, General Reward=✗2025.10 | 3 | — | |
| Foster et al. (2023b)Exploration Mechanism=information gain + optimism2025.10 | — | 2 |