Reinforcement Learning on Taxi environment stochastic
-57Episode Returnm-Trifle
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| m-TrifleSequence length K=7, Beam width N=8, Planning horizon H=32023.10 | -57 | 0.38 | 0.02 | |
| s-Trifle2023.10 | -99 | 0.14 | 0.11 | |
| dataset2023.10 | -128 | 2.41 | 0 | |
| DoC2023.10 | -146 | 0 | 0.28 | |
| TTSequence length K=7, Beam width N=8, Planning horizon H=32023.10 | -182 | 2.57 | 0.34 | |
| DTSequence length K=7, Conditioning RTG=-3002023.10 | -388 | 14.2 | 0.66 |