Reinforcement Learning on Maze
1.03Mean RewardProPS
Evaluation Results
| Method | Links | |
|---|---|---|
| ProPSNumber of runs=10, Aggregation protocol=averaged across all training iterations2026.05 | 1.03 | |
| ProPS+Number of independent runs=102026.05 | 0.97 | |
| A2CSource=Best SB3, Number of independent runs=102026.05 | 0.97 | |
| R2PONumber of independent runs=102026.05 | 0.97 | |
| A2CNumber of runs=10, Aggregation protocol=averaged across all training iterations, Source=Best SB32026.05 | 0.83 | |
| R2PONumber of runs=10, Aggregation protocol=averaged across all training iterations2026.05 | 0.83 | |
| ProPS+Number of runs=10, Aggregation protocol=averaged across all training iterations2026.05 | 0.76 | |
| ProPSNumber of independent runs=102026.05 | -0.7 |