Reinforcement Learning on LunarLander v3 (Average Agent Reward)
289Average Agent RewardSAC - H-EARS
Evaluation Results
| Method | Links | |
|---|---|---|
| SAC - H-EARSAlgorithm Variant=H-EARS, Base Algorithm=SAC2026.03 | 289 | |
| TD3 - VanillaAlgorithm Variant=Vanilla, Base Algorithm=TD32026.03 | 279 | |
| TD3 - H-EARSAlgorithm Variant=H-EARS, Base Algorithm=TD32026.03 | 277 | |
| SAC - VanillaAlgorithm Variant=Vanilla, Base Algorithm=SAC2026.03 | 268 | |
| PPO - H-EARSAlgorithm Variant=H-EARS, Base Algorithm=PPO2026.03 | 258 | |
| DDPG - H-EARSAlgorithm Variant=H-EARS, Base Algorithm=DDPG2026.03 | 250 | |
| POEMEvaluation Episodes=152026.01 | 242.1 | |
| PPO - VanillaAlgorithm Variant=Vanilla, Base Algorithm=PPO2026.03 | 235 | |
| DDPG - VanillaAlgorithm Variant=Vanilla, Base Algorithm=DDPG2026.03 | 231 | |
| PPOEvaluation Episodes=152026.01 | 210.94 | |
| PDAEnvironment steps=1M, Number of seeds=10, Tests per epoch=102026.03 | 204.7 | |
| PPOEnvironment steps=1M, Number of seeds=10, Tests per epoch=102026.03 | 204.4 | |
| NPGEnvironment steps=1M, Number of seeds=10, Tests per epoch=102026.03 | 34.1 | |
| TRPOEnvironment steps=1M, Number of seeds=10, Tests per epoch=102026.03 | -83 |