Reinforcement Learning on MountainCar v0
199.99Cumulative RewardDTSemNets
Evaluation Results
| Method | Links | |
|---|---|---|
| DTSemNetsAlgorithm=DTSemNets2026.05 | 199.99 | |
| πdisc.-PRLAlgorithm=π-PRL, Policy Stage=discretized policy before fine-tuning, Maximal Depth=62026.05 | 178.58 | |
| πcont.-PRLAlgorithm=π-PRL, Policy Stage=relaxed policy before discretization, Maximal Depth=62026.05 | 172.5 | |
| π-PRLAlgorithm=π-PRL, Policy Stage=final fine-tuned policy, Maximal Depth=62026.05 | 171.83 | |
| DiPRLAlgorithm=DiPRL, Maximal Depth=62026.05 | 110.79 | |
| PPOAlgorithm=PPO, Policy Type=Neural2026.05 | 98.57 | |
| VIPER (PPO)Algorithm=VIPER, Oracle Usage=PPO, Maximal Depth=62026.05 | 97.13 |