Regret Minimization on KL-regularized Bandits
2Sample ComplexityOnline Iterative GSHF
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Online Iterative GSHFCoverage=×2025.02 | 2 | — | — | |
| Two-Stage Mixed-Policy SamplingCoverage=✓2025.02 | 2 | — | — | |
| KL-UCBType=Upper Bound2026.03 | — | 2 | — |