POMDP Policy Optimization on Random POMDP
3.1284L^piCohen and Parmentier (2023)
Evaluation Results
| Method | Links | |
|---|---|---|
| Cohen and Parmentier (2023)|O| and |A|=32025.12 | 3.1284 | |
| Memoryless Policy Iteration|O| and |A|=22025.12 | 3.2125 | |
| Policy gradient|O| and |A|=22025.12 | 3.2125 | |
| Exhaustive search|O| and |A|=22025.12 | 3.2125 | |
| Cohen and Parmentier (2023)|O| and |A|=22025.12 | 3.2125 | |
| Müller and Montufar (2022)|O| and |A|=22025.12 | 3.2125 | |
| Müller and Montufar (2022)|O| and |A|=32025.12 | 3.2133 | |
| Memoryless Policy Iteration|O| and |A|=32025.12 | 3.214 | |
| Policy gradient|O| and |A|=32025.12 | 3.214 | |
| Exhaustive search|O| and |A|=32025.12 | 3.214 | |
| Müller and Montufar (2022)|O| and |A|=42025.12 | 3.3865 | |
| Memoryless Policy Iteration|O| and |A|=42025.12 | 3.3868 | |
| Policy gradient|O| and |A|=42025.12 | 3.3868 | |
| Müller and Montufar (2022)|O| and |A|=52025.12 | 3.4866 | |
| Memoryless Policy Iteration|O| and |A|=52025.12 | 3.492 | |
| Policy gradient|O| and |A|=52025.12 | 3.492 | |
| Müller and Montufar (2022)|O| and |A|=72025.12 | 3.5745 | |
| Memoryless Policy Iteration|O| and |A|=72025.12 | 3.5749 | |
| Policy gradient|O| and |A|=72025.12 | 3.5749 | |
| Müller and Montufar (2022)|O| and |A|=62025.12 | 3.6054 | |
| Policy gradient|O| and |A|=62025.12 | 3.6059 | |
| Memoryless Policy Iteration|O| and |A|=62025.12 | 3.606 | |
| Memoryless Policy Iteration|O| and |A|=82025.12 | 3.6338 | |
| Policy gradient|O| and |A|=82025.12 | 3.6338 | |
| Müller and Montufar (2022)|O| and |A|=82025.12 | 3.6338 | |
| Müller and Montufar (2022)|O| and |A|=102025.12 | 3.636 | |
| Müller and Montufar (2022)|O| and |A|=92025.12 | 3.6364 | |
| Memoryless Policy Iteration|O| and |A|=102025.12 | 3.6378 | |
| Policy gradient|O| and |A|=102025.12 | 3.6378 | |
| Memoryless Policy Iteration|O| and |A|=92025.12 | 3.6392 | |
| Policy gradient|O| and |A|=92025.12 | 3.6392 |