Constrained Reinforcement Learning on PointReach
89.2Episodic RewardCPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CPOCost Threshold=25, Number of Independent Runs=52025.12 | 89.2 | 33.3 | |
| e-COPCost Threshold=25, Number of Independent Runs=52025.12 | 81.5 | 24.5 | |
| P3OCost Threshold=25, Number of Independent Runs=52025.12 | 76.3 | 26.3 | |
| APPOCost Threshold=25, Number of Independent Runs=52025.12 | 74.3 | 26.3 | |
| PCPOCost Threshold=25, Number of Independent Runs=52025.12 | 73.2 | 24.9 | |
| FOCOPSCost Threshold=25, Number of Independent Runs=52025.12 | 65.1 | 24.8 | |
| IPOCost Threshold=25, Number of Independent Runs=52025.12 | 49.1 | 24.7 | |
| PPO-LCost Threshold=25, Number of Independent Runs=52025.12 | 46.1 | 25.1 |