Reinforcement Learning on RandomMDP
0.1207Average Update Time (s)SVI-SSP
Evaluation Results
| Method | Links | |
|---|---|---|
| SVI-SSPEpisodes=3000, H=15, iota=0.052021.06 | 0.1207 | |
| EB-SSPEpisodes=3000, iota=0.052021.06 | 0.2319 | |
| Bernstein-SSPEpisodes=3000, iota=22021.06 | 0.2918 | |
| Q-learning with e-greedyEpisodes=3000, epsilon=0.052021.06 | 0.3385 | |
| LCB-ADVANTAGE-SSPEpisodes=3000, H=5, iota=0.05, theta*=40962021.06 | 0.3517 | |
| UC-SSPEpisodes=3000, iota=12021.06 | 14.4472 | |
| ULCVIEpisodes=3000, H=80, iota=22021.06 | 15.7128 |