POMDP Planning on RockSample (15, 15)
20.53Expected ReturnAdaOPS
Evaluation Results
| Method | Links | |
|---|---|---|
| AdaOPSEvaluation Mode=Standard2025.10 | 20.53 | |
| AdaOPSEvaluation Protocol=per-size reference2025.10 | 20.53 | |
| GammaZeroEvaluation Mode=Full2025.10 | 20.5 | |
| BetaZeroEvaluation Mode=Full2025.10 | 19.87 | |
| DESPOTEvaluation Mode=Standard2025.10 | 18.83 | |
| DESPOTEvaluation Protocol=per-size reference2025.10 | 18.83 | |
| GammaZeroEvaluation Protocol=zero-shot transfer, Search Configuration=Full search2025.10 | 17.8 | |
| NVI|S|=7, 372, 800, |A|=20, |Z|=3, Number of runs=5, Simulations per run=10^52025.11 | 13.6 | |
| POMCGS|S|=7, 372, 800, |A|=20, |Z|=3, Number of runs=5, Simulations per run=10^52025.11 | 13.45 | |
| GammaZeroEvaluation Mode=Raw Pθ2025.10 | 11.1 | |
| GammaZeroEvaluation Protocol=zero-shot transfer, Search Configuration=Raw policy network (Pθ)2025.10 | 11.1 | |
| BetaZeroEvaluation Mode=Raw Pθ2025.10 | 11.04 | |
| POMCPOWEvaluation Mode=Standard2025.10 | 11.01 | |
| POMCPOWEvaluation Protocol=per-size reference2025.10 | 11.01 | |
| MCVI|S|=7, 372, 800, |A|=20, |Z|=3, Number of runs=5, Simulations per run=10^52025.11 | 10.78 | |
| BetaZeroEvaluation Mode=Raw V*θ2025.10 | 9.44 | |
| GammaZeroEvaluation Mode=Raw V*θ2025.10 | 9.1 | |
| GammaZeroEvaluation Protocol=zero-shot transfer, Search Configuration=One-step look-ahead value network (V*θ)2025.10 | 9.1 | |
| Recurrent PPO|S|=7, 372, 800, |A|=20, |Z|=3, Number of runs=5, Simulations per run=10^52025.11 | 6.16 |