Reinforcement Learning on CartPole v1 (test)
500Total RewardQualitatively measured policy discrepancy w/ β
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qualitatively measured policy discrepancy w/ βsource_paper=[23], instance=32020.12 | 500 | 1,157.2 | |
| Qualitatively measured policy discrepancy w/ βsource_paper=[23], instance=42020.12 | 500 | 1,157.2 | |
| Qualitatively measured policy discrepancy w/ ηsource_paper=[23], instance=42020.12 | 500 | 1,157.2 | |
| Differentiable Decision Treessource_paper=[3], note=Result confirmed by personal communication2020.12 | 500 | 106.8 | |
| Differentiable Decision Treessource_paper=[3], note=Tree simplified using the technique used in this work2020.12 | 500 | 53.4 | |
| two-level optimization schemeSplit Type=Orthogonal2020.12 | 500 | 35.6 | |
| two-level optimization schemeSplit Type=Oblique2020.12 | 500 | 21.1 | |
| Qualitatively measured policy discrepancy w/ βsource_paper=[23], instance=12020.12 | 499.9 | 1,157.2 | |
| General Q(λ)source_paper=[23]2020.12 | 499.9 | 1,157.2 | |
| Qualitatively measured policy discrepancy w/ ηsource_paper=[23], instance=32020.12 | 499.4 | 1,157.2 | |
| Importance-Samplingsource_paper=[23]2020.12 | 498.7 | 1,157.2 | |
| Peng & Williams's Q(λ)source_paper=[23]2020.12 | 496.7 | 1,157.2 | |
| Qualitatively measured policy discrepancy w/ βsource_paper=[23], instance=22020.12 | 494.9 | 1,157.2 | |
| Tree-Backup(λ)source_paper=[23]2020.12 | 494.7 | 1,157.2 | |
| Qualitatively measured policy discrepancy w/ ηsource_paper=[23], instance=22020.12 | 493.3 | 1,157.2 | |
| Qualitatively measured policy discrepancy w/ ηsource_paper=[23], instance=12020.12 | 493.2 | 1,157.2 | |
| Q(λ)source_paper=[23]2020.12 | 489.9 | 1,157.2 | |
| Watkins's Q(λ)source_paper=[23]2020.12 | 484.3 | 1,157.2 | |
| Retrace(λ)source_paper=[23]2020.12 | 461.1 | 1,157.2 | |
| Differentiable Decision Treessource_paper=[3]2020.12 | 388.76 | 89.2 | |
| Deep Q Networksource_paper=[23]2020.12 | 327.3 | 1,157.2 | |
| Kronecker-Factored Approximate Curvaturesource_paper=[25]2020.12 | 321 | 70,786.2 | |
| Bayesian Deep Reinforcement Learning weightedsource_paper=[24]2020.12 | 136.75 | 8,090.4 | |
| Bayesian Deep Reinforcement Learningsource_paper=[24]2020.12 | 113.52 | 8,090.4 | |
| Deep Q Networksource_paper=[24]2020.12 | 98.33 | 5,170,174.8 |