Reinforcement Learning on CartPole v1
354,122ReturnSat-EnQ
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Sat-EnQFunction Approximation=Neural Network (3-layer MLP), Weak learners=2-layer MLP (32 units), K=4, m=0.5, alpha=0.99, lambda=0.1, gamma=0.99, Training Steps=20,000, Seeds=102025.12 | 354,122 | 14,959 | 0 | 27 | — | |
| Maxmin Q-learningFunction Approximation=Neural Network (3-layer MLP), Training Steps=20,000, Seeds=102025.12 | 301,178 | 31,684 | 35 | 92 | — | |
| Double DQNFunction Approximation=Neural Network (3-layer MLP), Training Steps=20,000, Seeds=102025.12 | 287,205 | 42,025 | 40 | 43 | — | |
| Bootstrapped DQNFunction Approximation=Neural Network (3-layer MLP), Training Steps=20,000, Seeds=102025.12 | 268,195 | 38,025 | 40 | 105 | — | |
| DQNFunction Approximation=Neural Network (3-layer MLP), Training Steps=20,000, Seeds=102025.12 | 249,238 | 56,632 | 50 | 42 | — | |
| SAC-AdaGammaAdaptive-gamma=True2026.05 | 500 | — | — | — | — | |
| PPO-AdaGammaAdaptive-gamma=True2026.05 | 500 | — | — | — | — | |
| PPOAlgorithm=PPO, Policy Type=Neural2026.05 | 500 | — | — | — | — | |
| VIPER (PPO)Algorithm=VIPER, Oracle Usage=PPO, Maximal Depth=62026.05 | 500 | — | — | — | — | |
| πcont.-PRLAlgorithm=π-PRL, Policy Stage=relaxed policy before discretization, Maximal Depth=62026.05 | 500 | — | — | — | — | |
| πdisc.-PRLAlgorithm=π-PRL, Policy Stage=discretized policy before fine-tuning, Maximal Depth=62026.05 | 500 | — | — | — | — | |
| π-PRLAlgorithm=π-PRL, Policy Stage=final fine-tuned policy, Maximal Depth=62026.05 | 500 | — | — | — | — | |
| DiPRLAlgorithm=DiPRL, Maximal Depth=62026.05 | 500 | — | — | — | — | |
| SACAdaptive-gamma=False2026.05 | 483.2 | — | — | — | — | |
| PPOAdaptive-gamma=False2026.05 | 481.9 | — | — | — | — | |
| DTSemNetsAlgorithm=DTSemNets2026.05 | 326.76 | — | — | — | — | |
| Classical STC (B3)2026.05 | — | — | — | — | 0.212 | |
| DQNwc=162026.05 | — | — | — | — | 0.308 | |
| Pref-DQN (1 model)wc=142026.05 | — | — | — | — | 0.316 | |
| SACwc=162026.05 | — | — | — | — | 0.258 |