Trajectory Prediction on nuScenes (L2 Error and Collision Rate)
0.12L2 Error (m) 1sDriveTeach-VLA-3B
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| DriveTeach-VLA-3BModel Category=Specialized Driving Models, Model Size=3B2026.07 | 0.12 | 0.27 | 0.51 | 0.3 | — | — | — | — | |
| Impromptu-VLA-3BModel Category=Specialized Driving Models, Model Size=3B2026.07 | 0.13 | 0.27 | 0.52 | 0.31 | — | — | — | — | |
| Impromptu-VLA-7BModel Category=Specialized Driving Models, Model Size=7B2026.07 | 0.13 | 0.27 | 0.53 | 0.31 | — | — | — | — | |
| OmniDriveModel Category=Specialized Driving Models2026.07 | 0.14 | 0.29 | 0.55 | 0.33 | — | — | — | — | |
| EMMAModel Category=Specialized Driving Models2026.07 | 0.14 | 0.29 | 0.54 | 0.32 | — | — | — | — | |
| Ego-MLP*Model Category=Training-based Driving Specialists (Existing Methods)2026.07 | 0.15 | 0.32 | 0.59 | 0.35 | — | — | — | — | |
| DriveVLM-DualModel Category=Specialized Driving Models2026.07 | 0.15 | 0.29 | 0.48 | 0.31 | — | — | — | — | |
| EMMA (random init)Model Category=Specialized Driving Models, initialization=random init2026.07 | 0.15 | 0.33 | 0.63 | 0.37 | — | — | — | — | |
| BEV-PlannerModel Category=Training-based Driving Specialists (Existing Methods)2026.07 | 0.16 | 0.32 | 0.57 | 0.35 | — | — | — | — | |
| VADModel Category=Training-based Driving Specialists (Existing Methods)2026.07 | 0.17 | 0.34 | 0.6 | 0.37 | — | — | — | — | |
| DriveVLMModel Category=Specialized Driving Models2026.07 | 0.18 | 0.34 | 0.68 | 0.4 | — | — | — | — | |
| PRIXInput=Camera, Backbone=ResNet-50, FPS=11.22025.07 | 0.26 | 0.53 | 0.93 | 0.57 | 0 | 4 | 0.18 | 7 | |
| GPT-4oModel Category=Generalist VLMs2026.07 | 0.28 | 0.93 | 2.02 | 1.07 | — | — | — | — | |
| Claude-3.7-SonnetModel Category=Generalist VLMs2026.07 | 0.28 | 0.94 | 2.04 | 1.09 | — | — | — | — | |
| SparseDriveInput=Camera, Backbone=ResNet-50, FPS=9.02025.07 | 0.29 | 0.58 | 0.96 | 0.61 | 1 | 5 | 0.18 | 8 | |
| DiffusionDriveInput=Camera, Backbone=ResNet-50, FPS=8.22025.07 | 0.31 | 0.62 | 1.03 | 0.65 | 3 | 6 | 0.19 | 9 | |
| Gemini-2.5-ProModel Category=Generalist VLMs2026.07 | 0.37 | 1.35 | 2.96 | 1.56 | — | — | — | — | |
| VADInput=Camera, Backbone=ResNet-50, FPS=4.52025.07 | 0.41 | 0.7 | 1.05 | 0.72 | 7 | 17 | 0.41 | 22 | |
| UniADModel Category=Training-based Driving Specialists (Existing Methods)2026.07 | 0.42 | 0.64 | 0.91 | 0.66 | — | — | — | — | |
| UniADInput=Camera, Backbone=ResNet-101, FPS=1.82025.07 | 0.45 | 0.7 | 1.04 | 0.73 | 62 | 58 | 0.63 | 61 | |
| Qwen-2.5-VL-7B-InstructModel Category=Generalist VLMs, Model Size=7B2026.07 | 0.46 | 1.33 | 2.55 | 1.45 | — | — | — | — | |
| LLaMA-3.2-11B-Vision-InstructModel Category=Generalist VLMs, Model Size=11B2026.07 | 0.52 | 1.42 | 2.68 | 1.54 | — | — | — | — | |
| DeepSeek-VL2-16BModel Category=Generalist VLMs, Model Size=16B2026.07 | 0.66 | 1.68 | 2.92 | 1.75 | — | — | — | — | |
| OccNetInput=Camera, Backbone=ResNet-50, FPS=2.62025.07 | 1.29 | 2.13 | 2.99 | 2.14 | 21 | 59 | 1.37 | 72 | |
| ST-P3Input=Camera, Backbone=EffNet-b4, FPS=1.62025.07 | 1.33 | 2.11 | 2.9 | 2.11 | 23 | 62 | 1.27 | 71 | |
| Qwen2-VL-7B-InstructModel Category=Generalist VLMs, Model Size=7B2026.07 | 1.45 | 3.21 | 3.76 | 2.81 | — | — | — | — | |
| LLaVA-1.6-Mistral-7BModel Category=Generalist VLMs, Model Size=7B2026.07 | 1.49 | 3.38 | 4.09 | 2.98 | — | — | — | — |