Decision Making under Imperfect Recall on 61 aggregate benchmark instances (Sim, Det, and Rand)
73.8Utility (% of Best)PRM
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| PRM2026.02 | 73.8 | 3.79 | 78.7 | 18 | 4.24 | |
| RM2026.02 | 70.5 | 3.93 | 83.6 | 29.5 | 2.8 | |
| RM+2026.02 | 70.5 | 3.79 | 90.2 | 21.3 | 2.7 | |
| PRM+2026.02 | 70.5 | 3.81 | 80.3 | 18 | 3.99 | |
| AMShyperparameter combinations=27, learning rates=η ∈ {1, 0.1, 0.01}, beta coefficients=β1 ∈ {0.8, 0.9, 0.99}, β2 ∈ {0.99, 0.999, 0.9999}2026.02 | 60.7 | 4.16 | 98.4 | 37.7 | 3.35 | |
| GDlearning rates=η ∈ {1, 10^-1, 10^-2, 10^-3}2026.02 | 44.3 | 4.99 | 72.1 | 11.5 | 5.08 | |
| OGDlearning rates=η ∈ {1, 10^-1, 10^-2, 10^-3}2026.02 | 44.3 | 5.28 | 72.1 | 6.6 | 5.84 | |
| Gurobi2026.02 | 36.1 | 6.25 | 27.9 | 0 | 8 |