Reinforcement Learning on InvertedPendulum
1,000Mean RewardR2PO
Evaluation Results
| Method | Links | |
|---|---|---|
| R2PONumber of independent runs=102026.05 | 1,000 | |
| R2PONumber of runs=10, Aggregation protocol=averaged across all training iterations2026.05 | 756.08 | |
| ProPS+Number of independent runs=102026.05 | 657.88 | |
| ProPSNumber of independent runs=102026.05 | 649.91 | |
| ProPS+Number of runs=10, Aggregation protocol=averaged across all training iterations2026.05 | 309.52 | |
| ProPSNumber of runs=10, Aggregation protocol=averaged across all training iterations2026.05 | 234.14 | |
| TRPOSource=Best SB3, Number of independent runs=102026.05 | 28.49 | |
| TRPONumber of runs=10, Aggregation protocol=averaged across all training iterations, Source=Best SB32026.05 | 24.35 |