Safe Reinforcement Learning on FormulaOne (L1)
64.1JRVLMPPOLag
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| VLMPPOLagType=Ours (CMDP+VLM cost), Seeds=32026.06 | 64.1 | 32.8 | |
| PPOLag-DecoupledType=Ours (CMDP+VLM cost), Seeds=32026.06 | 64 | 33.5 | |
| CPO-DecoupledType=Ours (CMDP+VLM cost), Seeds=32026.06 | 63.7 | 37.6 | |
| PPO-CLGType=CLG (VLM-as-reward), Seeds=32026.06 | 51.7 | 133.6 | |
| CPO-CLGType=CLG (VLM-as-reward), Seeds=32026.06 | 50.7 | 32.8 | |
| VLMPPOLag+Conf*Type=Ours (CMDP+VLM cost), Seeds=5, Gating=Calibrated gate from Eq. (5)2026.06 | 33.5 | 20.8 | |
| CPO-CoupledType=Ours (CMDP+VLM cost), Seeds=32026.06 | 21 | 29.1 | |
| PPOType=RL (pure RL), Seeds=32026.06 | 1.6 | 217 | |
| PPOLagType=CMDP (no VLM), Seeds=32026.06 | 0.8 | 67.9 | |
| PPOLag-RNDType=RND (intrinsic-novelty ablation), Seeds=32026.06 | 0.8 | 62.4 | |
| CPOType=CMDP (no VLM), Seeds=32026.06 | 0.3 | 35.7 |