Safe Reinforcement Learning on FormulaOne (L0)
64.3JRPPOLag-Decoupled
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| PPOLag-DecoupledType=Ours (CMDP+VLM cost), Seeds=32026.06 | 64.3 | 0 | |
| VLMPPOLagType=Ours (CMDP+VLM cost), Seeds=32026.06 | 64.3 | 0 | |
| CPO-DecoupledType=Ours (CMDP+VLM cost), Seeds=32026.06 | 64.1 | 0 | |
| PPO-CLGType=CLG (VLM-as-reward), Seeds=32026.06 | 51.9 | 0 | |
| CPO-CLGType=CLG (VLM-as-reward), Seeds=32026.06 | 51.7 | 0 | |
| VLMPPOLag+Conf*Type=Ours (CMDP+VLM cost), Seeds=5, Gating=Calibrated gate from Eq. (5)2026.06 | 44.5 | 0 | |
| CPO-CoupledType=Ours (CMDP+VLM cost), Seeds=32026.06 | 21.2 | 0 | |
| CPOType=CMDP (no VLM), Seeds=32026.06 | 1.9 | 0 | |
| PPOType=RL (pure RL), Seeds=32026.06 | 1.6 | 0 | |
| PPOLagType=CMDP (no VLM), Seeds=32026.06 | 1.6 | 0 |