Constrained Reinforcement Learning on Navigation
217.6Episodic Rewarde-COP
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| e-COPCost Threshold C1=10, Cost Threshold C2=25, Number of Independent Runs=52025.12 | 217.6 | 9.6 | 23.7 | |
| PPO-LCost Threshold C1=10, Cost Threshold C2=25, Number of Independent Runs=52025.12 | 175.1 | 9.9 | 22.3 | |
| IPOCost Threshold C1=10, Cost Threshold C2=25, Number of Independent Runs=52025.12 | 164.1 | 10 | 24.6 | |
| P3OCost Threshold C1=10, Cost Threshold C2=25, Number of Independent Runs=52025.12 | 153.5 | 9.9 | 24.5 | |
| APPOCost Threshold C1=10, Cost Threshold C2=25, Number of Independent Runs=52025.12 | 135.7 | 9.9 | 23.9 |