Explanation Generation on VecEval (val)
93.26B ScoreFleetAgent
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| FleetAgentInput Modality=Vectorized V2N Messages2026.06 | 93.26 | 33.24 | 27.52 | |
| Gemini-2.5-Flash (Reference Baseline)Input Modality=Language Description2026.06 | 89.33 | 38.86 | 37.53 | |
| Qwen-2.5VL-7BInput Modality=Raw Images, Evaluation Protocol=Few-shot Example2026.06 | 66.28 | 26.43 | 27.58 | |
| GPT-4oInput Modality=BEV Images2026.06 | 62.9 | 24.16 | 20.92 | |
| FleetAgent (w/o tokens prioritization)Input Modality=Vectorized V2N Messages2026.06 | 61.99 | 36.58 | 31.1 | |
| GPT-4oInput Modality=Raw Images2026.06 | 61.65 | 24.59 | 22.34 | |
| Qwen-2.5VL-7BInput Modality=Language Description, Evaluation Protocol=Few-shot Example2026.06 | 48.38 | 29.84 | 22.21 | |
| GPT-4oInput Modality=Language Description2026.06 | 45.05 | 29.07 | 22.21 | |
| Qwen-2.5VL-7BInput Modality=BEV Images, Evaluation Protocol=Few-shot Example2026.06 | 42.16 | 21.84 | 17.78 |