POMDP Planning on RockSample (20, 20)
12.31Expected ReturnNVI
Evaluation Results
| Method | Links | |
|---|---|---|
| NVI|S|=419, 430, 400, |A|=25, |Z|=3, Number of runs=5, Simulations per run=10^52025.11 | 12.31 | |
| AdaOPSEvaluation Protocol=per-size reference2025.10 | 10.96 | |
| GammaZeroEvaluation Protocol=zero-shot transfer, Search Configuration=Full search2025.10 | 10.2 | |
| POMCPOWEvaluation Protocol=per-size reference2025.10 | 9.92 | |
| GammaZeroEvaluation Protocol=zero-shot transfer, Search Configuration=Raw policy network (Pθ)2025.10 | 5.4 | |
| MCVI|S|=419, 430, 400, |A|=25, |Z|=3, Number of runs=5, Simulations per run=10^52025.11 | 4.57 | |
| GammaZeroEvaluation Protocol=zero-shot transfer, Search Configuration=One-step look-ahead value network (V*θ)2025.10 | 4.4 | |
| Recurrent PPO|S|=419, 430, 400, |A|=25, |Z|=3, Number of runs=5, Simulations per run=10^52025.11 | 3.97 | |
| POMCGS|S|=419, 430, 400, |A|=25, |Z|=3, Number of runs=5, Simulations per run=10^52025.11 | 3.91 | |
| DESPOTEvaluation Protocol=per-size reference2025.10 | 0 |