Constrained Reinforcement Learning on Bottleneck
388.1Episodic RewardCPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CPOCost Threshold=50, Number of Independent Runs=52025.12 | 388.1 | 54.3 | |
| e-COPCost Threshold=50, Number of Independent Runs=52025.12 | 345.1 | 49.7 | |
| PPO-LCost Threshold=50, Number of Independent Runs=52025.12 | 298.3 | 41.4 | |
| P3OCost Threshold=50, Number of Independent Runs=52025.12 | 291.1 | 45.3 | |
| IPOCost Threshold=50, Number of Independent Runs=52025.12 | 279.3 | 48.2 | |
| PCPOCost Threshold=50, Number of Independent Runs=52025.12 | 264.2 | 49.8 | |
| FOCOPSCost Threshold=50, Number of Independent Runs=52025.12 | 251.3 | 46.6 | |
| APPOCost Threshold=50, Number of Independent Runs=52025.12 | 220.1 | 47.4 |