Reinforcement Learning on Chain (Sample Complexity < 0.4 Regret)
61Sample ComplexityRandom Exploration
Evaluation Results
| Method | Links | |
|---|---|---|
| Random ExplorationNE=12022.07 | 61 | |
| TRAVEL (gener. model)NE=12022.07 | 76 | |
| Uniform sampling (gener. model)NE=12022.07 | 78 | |
| AceIRL (Full)NE=12022.07 | 142 | |
| AceIRL GreedyNE=12022.07 | 153 |