Reinforcement Learning on DoorKey
0.97RewardPPO
Evaluation Results
| Method | Links | |
|---|---|---|
| PPOAlgorithm=PPO, Policy Type=Neural2026.05 | 0.97 | |
| DiPRLAlgorithm=DiPRL, Maximal Depth=62026.05 | 0.95 | |
| π-PRLAlgorithm=π-PRL, Policy Stage=final fine-tuned policy, Maximal Depth=62026.05 | 0.66 | |
| πcont.-PRLAlgorithm=π-PRL, Policy Stage=relaxed policy before discretization, Maximal Depth=62026.05 | 0.63 | |
| VIPER (PPO)Algorithm=VIPER, Oracle Usage=PPO, Maximal Depth=62026.05 | 0.55 | |
| πdisc.-PRLAlgorithm=π-PRL, Policy Stage=discretized policy before fine-tuning, Maximal Depth=62026.05 | 0.32 | |
| DTSemNetsAlgorithm=DTSemNets2026.05 | 0 |