Reinforcement Learning on Procgen easy levels zero-shot generalization (test)
0.2969bigfishVPN
Evaluation Results
| Method | Links | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| VPNMethod Category=RL+value pred., Generalization=Zero-shot, Seeds=52021.11 | 0.2969 | -0.7571 | -0.686 | -0.849 | -1 | -0.4973 | 0.0102 | -0.7804 | -0.3623 | -0.1961 | 0.4571 | -0.2456 | -0.3939 | -0.6443 | -0.7843 | -0.75 | |
| PSEMethod Category=RL+GVF dist., Generalization=Zero-shot, Seeds=52021.11 | 0.1898 | -0.3557 | -0.2215 | -0.349 | -0.2632 | 0.2869 | -0.1122 | -0.2626 | 0.0348 | 0.0353 | 0.3749 | 0.0211 | 0.1273 | -0.1537 | -0.3247 | -0.2207 | |
| GSF (reward)Method Category=RL+GVF dist., Generalization=Zero-shot, Seeds=52021.11 | 0.1536 | 0.3857 | 0.124 | -0.0495 | 1.8947 | 0.4664 | -0.0306 | 0.2383 | -0.1594 | 0.8627 | -0.0914 | 0.1228 | 0.1212 | 0.0448 | -0.0784 | 0.1782 | |
| CSSCMethod Category=RL+action dist., Generalization=Zero-shot, Seeds=52021.11 | 0.116 | -0.8786 | -0.6694 | -0.9191 | -0.8596 | -0.5972 | -0.102 | -0.7477 | -0.2754 | -0.3333 | 0.4971 | -0.2281 | -0.4242 | -0.6667 | -0.7843 | -0.7672 | |
| DeepMDPMethod Category=RL+Bisim., Generalization=Zero-shot, Seeds=52021.11 | -0.2969 | -0.0331 | -0.7273 | -0.8902 | -0.5965 | -0.5972 | -0.2449 | -0.7477 | -0.4203 | -0.549 | 0.0743 | -0.2544 | -0.4545 | -0.8682 | -0.7686 | -0.773 | |
| BCMethod Category=RL+Bisim., Generalization=Zero-shot, Seeds=52021.11 | -0.4437 | 0.319 | -0.9452 | 0.1484 | 0.6316 | -0.742 | -0.1327 | -0.3738 | -0.6377 | 0.1961 | -0.2857 | -0.3684 | -0.0606 | -0.097 | 0.1569 | -0.023 |