Safe Reinforcement Learning on MetaDrive Hard (held-out seeds 10000–10019)
25Categorical Violation Rate (%)VLMPPOLag+Conf
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| VLMPPOLag+Confwarm-start λ0=0.52026.06 | 25 | 28 | 26 | — | |
| VLMPPOLag+Confwarm-start λ0=default2026.06 | 31 | 39 | — | — | |
| PPOLagwarm-start λ0=default2026.06 | 33 | 36 | — | — | |
| PPOLagwarm-start λ0=0.52026.06 | 33 | 36 | — | — |