Safe Reinforcement Learning on FormulaOne (L2)
63.9JR ScoreCPO-Decoupled
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CPO-DecoupledType=Ours (CMDP+VLM cost), Seeds=32026.06 | 63.9 | 30.9 | |
| PPOLag-DecoupledType=Ours (CMDP+VLM cost), Seeds=32026.06 | 63.8 | 40.7 | |
| VLMPPOLagType=Ours (CMDP+VLM cost), Seeds=32026.06 | 63.8 | 40.2 | |
| PPO-CLGType=CLG (VLM-as-reward), Seeds=32026.06 | 51.3 | 156.6 | |
| CPO-CLGType=CLG (VLM-as-reward), Seeds=32026.06 | 50.9 | 33.9 | |
| VLMPPOLag+Conf*Type=Ours (CMDP+VLM cost), Seeds=5, Gating=Calibrated gate from Eq. (5)2026.06 | 31.8 | 22.5 | |
| CPO-CoupledType=Ours (CMDP+VLM cost), Seeds=32026.06 | 21.6 | 32.4 | |
| PPOType=RL (pure RL), Seeds=32026.06 | 1.3 | 269.2 | |
| PPOLagType=CMDP (no VLM), Seeds=32026.06 | 0.7 | 55.8 | |
| PPOLag-RNDType=RND (intrinsic-novelty ablation), Seeds=32026.06 | 0.4 | 45.7 | |
| CPOType=CMDP (no VLM), Seeds=32026.06 | 0.3 | 36.1 | |
| CPPOPIDType=CMDP (no VLM), Seeds=32026.06 | 0.2 | 22.8 |