Reinforcement Learning on Inverted Double Pendulum
9,359.92Avg Episode RewardSAC
Evaluation Results
| Method | Links | |
|---|---|---|
| SAC2023.11 | 9,359.92 | |
| ESPL2023.11 | 9,359.9 | |
| A2C2023.11 | 9,359.81 | |
| TD32023.11 | 9,359.25 | |
| ACKTR2023.11 | 9,359.06 | |
| PPO2023.11 | 9,356.59 | |
| DDPG2023.11 | 9,347.1 | |
| TRPO2023.11 | 9,188.43 | |
| DSP2023.11 | 9,149.9 | |
| Regression2023.11 | 637.2 | |
| R2PONumber of independent runs=102026.05 | 254.04 | |
| R2PONumber of runs=10, Aggregation protocol=averaged across all training iterations2026.05 | 158.51 | |
| ProPS+Number of independent runs=102026.05 | 128.81 | |
| ProPSNumber of independent runs=102026.05 | 112.18 | |
| TRPOSource=Best SB3, Number of independent runs=102026.05 | 98.25 | |
| ProPS+Number of runs=10, Aggregation protocol=averaged across all training iterations2026.05 | 86.71 | |
| TRPONumber of runs=10, Aggregation protocol=averaged across all training iterations, Source=Best SB32026.05 | 86.04 | |
| ProPSNumber of runs=10, Aggregation protocol=averaged across all training iterations2026.05 | 79.44 |