Closed-loop Planning on Bench2Drive zero-shot
57.1Driving ScoreRL (Rule-PDMS + RM CF&LG)
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| RL (Rule-PDMS + RM CF&LG)Reward formulation=composite reward (rule-based PDMS + CF and LG scores predicted by DriveReward-1B), Optimization=Reinforcement Learning2026.06 | 57.1 | 27.2 | 128.3 | 11.4 | |
| RL (Rule-Based PDMS)Reward formulation=standard rule-based PDMS, Optimization=Reinforcement Learning2026.06 | 55 | 25 | 124.4 | 14.8 | |
| RL (RM-Predicted PDMS)Reward formulation=PDMS directly predicted by DriveReward-1B, Optimization=Reinforcement Learning2026.06 | 51.4 | 20.8 | 124 | 19.2 | |
| UniAD-B.2026.06 | 45.8 | 16.4 | 129.2 | 43.6 | |
| VAD2026.06 | 42.4 | 15 | 157.9 | 46 | |
| UniAD-T.2026.06 | 40.7 | 13.2 | 123.9 | 47 | |
| Base SFTPolicy Model=InternVL3-8B, Training=Supervised Fine-tuning2026.06 | 40.2 | 14.5 | 122.1 | 42.9 | |
| AD-MLP2026.06 | 18.1 | 0 | 48.5 | 22.6 |