Constrained Reinforcement Learning on Grid
276.3Episodic RewardPPO-L
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| PPO-LCost Threshold=75, Number of Independent Runs=52025.12 | 276.3 | 71.8 | |
| e-COPCost Threshold=75, Number of Independent Runs=52025.12 | 258.1 | 71.3 | |
| IPOCost Threshold=75, Number of Independent Runs=52025.12 | 229.4 | 74.2 | |
| PCPOCost Threshold=75, Number of Independent Runs=52025.12 | 226.5 | 72.6 | |
| FOCOPSCost Threshold=75, Number of Independent Runs=52025.12 | 215.4 | 76.6 | |
| P3OCost Threshold=75, Number of Independent Runs=52025.12 | 201.5 | 79.3 | |
| APPOCost Threshold=75, Number of Independent Runs=52025.12 | 184.4 | 79.5 | |
| CPOCost Threshold=75, Number of Independent Runs=52025.12 | 178.1 | 69.3 |