Safe Reinforcement Learning on 15-concept neural simulation environment (test)
34.64ReturnUnconstrained
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Unconstrained2026.04 | 34.64 | 69.92 | 91.7 | 19.02 | 27.14 | |
| Reward-Shaped2026.04 | 34.64 | 69.92 | 91.7 | 19.02 | 27.14 | |
| Post-hoc2026.04 | 34.64 | 69.92 | 91.7 | 19.02 | 27.14 | |
| MC-CPOfrontier=enabled2026.04 | 32.73 | 44.47 | 87.1 | 9.69 | 23.14 | |
| MC-CPO (no frontier)frontier=disabled2026.04 | 32.68 | 44.51 | 87.2 | 9.69 | 23.12 |