Constrained Reinforcement Learning on PointCircle
110.5Episodic Rewarde-COP
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| e-COPCost Threshold=10, Number of Independent Runs=52025.12 | 110.5 | 9.8 | |
| APPOCost Threshold=10, Number of Independent Runs=52025.12 | 91.2 | 10.2 | |
| P3OCost Threshold=10, Number of Independent Runs=52025.12 | 89.1 | 9.9 | |
| FOCOPSCost Threshold=10, Number of Independent Runs=52025.12 | 81.6 | 10 | |
| IPOCost Threshold=10, Number of Independent Runs=52025.12 | 68.7 | 9.3 | |
| PCPOCost Threshold=10, Number of Independent Runs=52025.12 | 68.2 | 9.9 | |
| CPOCost Threshold=10, Number of Independent Runs=52025.12 | 65.3 | 9.5 | |
| PPO-LCost Threshold=10, Number of Independent Runs=52025.12 | 57.2 | 9.8 |