Reinforcement Learning on LunarLander v2
2,292Final ReturnAdvantage-weighting
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Advantage-weightingSource=Peng et al.2020.12 | 2,292 | 518,153 | |
| PPOAlgorithm=PPO, Policy Type=Neural2026.05 | 283.11 | — | |
| Deep Q NetworkSource=Xinly Yu2020.12 | 278.23 | 518,153 | |
| Oblique DTCriteria=Best Score2020.12 | 272.14 | 118.9 | |
| Deep Q NetworkSource=Sigve Rokenes2020.12 | 266 | 121,000,000 | |
| Oblique DTCriteria=Best M2020.12 | 262.18 | 86.9 | |
| DiPRLAlgorithm=DiPRL, Maximal Depth=62026.05 | 260.21 | — | |
| Shallow NNSource=Malagon et al.2020.12 | 258.8 | 77.6 | |
| DTSemNetsAlgorithm=DTSemNets2026.05 | 257.61 | — | |
| π-PRLAlgorithm=π-PRL, Policy Stage=final fine-tuned policy, Maximal Depth=62026.05 | 257.21 | — | |
| Actor CriticSource=Nikhil Barhate2020.12 | 254.58 | 4,337.3 | |
| Value-differenceSource=Xu et al.2020.12 | 248.221 | 632,620.2 | |
| πdisc.-PRLAlgorithm=π-PRL, Policy Stage=discretized policy before fine-tuning, Maximal Depth=62026.05 | 239.9 | — | |
| πcont.-PRLAlgorithm=π-PRL, Policy Stage=relaxed policy before discretization, Maximal Depth=62026.05 | 235.72 | — | |
| AWR2019.10 | 229 | — | |
| Deep Q NetworkSource=Ash Bellet2020.12 | 225.79 | 1,295,307.1 | |
| Soft Actor CriticSource=Keavan2020.12 | 217.92 | 210,733.2 | |
| Soft Q NetworkSource=liu2020.12 | 217.09 | 647,691.1 | |
| Proximal Policy Opt.Source=Daniel Barbosa2020.12 | 201.47 | 1,673 | |
| Deep Q NetworkSource=Ollie Graham2020.12 | 201.46 | 30,878.1 | |
| Deep Q NetworkSource=Udacity2020.12 | 201.46 | 30,878.1 | |
| Deep Q NetworkSource=Sanket Thakur2020.12 | 200.65 | 259,285.8 | |
| Deep Q NetworkSource=Mahmood2020.12 | 200.3 | 237,079.7 | |
| Dueling Deep Q N.Source=Ruslan2020.12 | 200.22 | 30,878.1 | |
| RWR2019.10 | 185 | — | |
| VIPER (PPO)Algorithm=VIPER, Oracle Usage=PPO, Maximal Depth=62026.05 | 164.33 | — | |
| NLDT*Source=Dhebar et al., Depth=32020.12 | 132.83 | 136.7 | |
| PPO2019.10 | 121 | — | |
| TRPO2019.10 | 104 | — | |
| Rule ListSource=Silva et al.2020.12 | -78.4 | 89 | |
| Oblique DTCriteria=Mean2020.12 | — | 123.3439 |