Reinforcement Learning on Tabular MDP
0.5Sample ComplexityLog-barrier, Momentum
Evaluation Results
| Method | Links | |
|---|---|---|
| Log-barrier, MomentumLearning Rate α=O(1/√t), Assumption on c*†=✗2026.03 | 0.5 |
| Method | Links | |
|---|---|---|
| Log-barrier, MomentumLearning Rate α=O(1/√t), Assumption on c*†=✗2026.03 | 0.5 |