Reinforcement Learning on MountainCar v0 (test)
-101.72Total RewardOrthogonal DT
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Orthogonal DTOptimization Criterion=Best Score2020.12 | -101.72 | 106.8 | |
| Closed-form policySource=Zhiqing Xiao2020.12 | -102.61 | 54.7 | |
| Soft Q NetworksSource=Keavnn2020.12 | -104.58 | 31,079.2 | |
| Tabular SARSASource=Amit2020.12 | -105.99 | 381.5 | |
| Oblique DTOptimization Criterion=Best Score2020.12 | -106.02 | 46.8 | |
| Double Deep Q NetworkSource=Colin M2020.12 | -107.83 | 46,681.6 | |
| Deep Q NetworkSource=Harshit Singh2020.12 | -108.85 | 984,160.3 | |
| Orthogonal DTOptimization Criterion=Best M2020.12 | -116.68 | 35.6 | |
| Nonlinear DT (Open loop)Source=Dhebar et al.2020.12 | -128.87 | 66.8 | |
| Oblique DTOptimization Criterion=Best M2020.12 | -200 | 0 |