POMDP Simulation on RS (7, 8, 20, 0)
28.5RewardNaive
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Naivev=12022.09 | 28.5 | 15.3 | |
| Perfect2022.09 | 28.4 | — | |
| Noisy Agentlambda=22022.09 | 27.8 | 9.1 | |
| Noisy Agentlambda=52022.09 | 27.5 | 7.8 | |
| Scaled Agenttau=0.992022.09 | 27.4 | 6.4 | |
| Scaled Agenttau=0.752022.09 | 27.3 | 6.8 | |
| Scaled Agenttau=0.52022.09 | 27 | 7.8 | |
| Noisy Agentlambda=12022.09 | 26.8 | 10.6 | |
| Naivev=0.752022.09 | 26 | 15.3 | |
| Naivev=0.52022.09 | 23.8 | 15.1 | |
| Normal2022.09 | 21.5 | — |