Frame-wise Semantic Risk Detection on CARLA Long-tail Semantic Failures (in-distribution)
91.8Emergency Vehicle Accuracy (Cat 1)VLM-Only
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| VLM-Onlybackbone=GPT-5-mini2025.12 | 91.8 | 84.07 | 93.49 | 90.36 | 97.14 | 96.71 | 94.14 | 90.38 | |
| LSREtype=Lightweight classifier distilled from VLM2025.12 | 85.66 | 98.64 | 89.54 | 91.61 | 93.32 | 93.08 | 89.51 | 94.44 | |
| Always-Safebehavior=Always predicts safe2025.12 | 54.36 | 0 | 33.49 | 0 | 27.6 | 0 | 38.48 | 0 |