Frame-wise Semantic Risk Detection on CARLA Long-tail Semantic Failures Few-shot
93.38Category 1: Emergency Vehicle AccuracyVLM-Only
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| VLM-Onlybackbone=GPT-5-mini2025.12 | 93.38 | 98.53 | 79.4 | 74.36 | 81.69 | 96.24 | 84.82 | 89.71 | |
| LSREtype=Lightweight classifier distilled from VLM2025.12 | 91.29 | 92.06 | 72.57 | 76.9 | 79.88 | 93.14 | 81.25 | 87.37 | |
| Always-Safebehavior=Always predicts safe2025.12 | 66.64 | 0 | 42.03 | 0 | 59.79 | 0 | 56.15 | 0 |