Reinforcement Learning on GridWorld
0.1419Average Update Time (s)SVI-SSP
Evaluation Results
| Method | Links | |
|---|---|---|
| SVI-SSPEpisodes=3000, H=10, iota=0.012021.06 | 0.1419 | |
| Q-learning with e-greedyEpisodes=3000, epsilon=0.052021.06 | 0.3773 | |
| LCB-ADVANTAGE-SSPEpisodes=3000, H=5, iota=0.1, theta*=40962021.06 | 0.3982 | |
| EB-SSPEpisodes=3000, iota=0.012021.06 | 0.4619 | |
| Bernstein-SSPEpisodes=3000, iota=0.52021.06 | 0.4656 | |
| UC-SSPEpisodes=3000, iota=0.52021.06 | 8.6886 | |
| ULCVIEpisodes=3000, H=100, iota=12021.06 | 22.8062 |