Constrained Reinforcement Learning on AntReach
102.3Episodic RewardCPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CPOCost Threshold=25, Number of Independent Runs=52025.12 | 102.3 | 35.1 | |
| P3OCost Threshold=25, Number of Independent Runs=52025.12 | 73.6 | 24.8 | |
| e-COPCost Threshold=25, Number of Independent Runs=52025.12 | 70.8 | 24.2 | |
| APPOCost Threshold=25, Number of Independent Runs=52025.12 | 61.5 | 24.5 | |
| PPO-LCost Threshold=25, Number of Independent Runs=52025.12 | 54.2 | 21.9 | |
| FOCOPSCost Threshold=25, Number of Independent Runs=52025.12 | 48.3 | 25.1 | |
| IPOCost Threshold=25, Number of Independent Runs=52025.12 | 45.2 | 24.9 | |
| PCPOCost Threshold=25, Number of Independent Runs=52025.12 | 39.4 | 27.9 |