Reinforcement Learning on Tabular CMDP (last 1,000 episodes)
0.999ReturnUnconstrained
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Unconstrained2026.04 | 0.999 | 99.9 | |
| MC-CPO2026.04 | 0.6 | 0.04 | |
| Post-hoc2026.04 | 0.0005 | 99.9 |
| Method | Links | ||
|---|---|---|---|
| Unconstrained2026.04 | 0.999 | 99.9 | |
| MC-CPO2026.04 | 0.6 | 0.04 | |
| Post-hoc2026.04 | 0.0005 | 99.9 |