Context-aware Evaluation on VecEval (val)
78.61Lingo-Judge AccuracyGemini-2.5-Flash (Reference Baseline)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Gemini-2.5-Flash (Reference Baseline)Input Modality=Language Description2026.06 | 78.61 | 44.78 | |
| FleetAgentInput Modality=Vectorized V2N Messages2026.06 | 55.93 | 35.68 | |
| Qwen-2.5VL-7BInput Modality=Raw Images, Evaluation Protocol=Few-shot Example2026.06 | 55.63 | 32.81 | |
| FleetAgent (w/o tokens prioritization)Input Modality=Vectorized V2N Messages2026.06 | 54.44 | 35.99 | |
| Qwen-2.5VL-7BInput Modality=Language Description, Evaluation Protocol=Few-shot Example2026.06 | 52.22 | 30.56 | |
| GPT-4oInput Modality=Language Description2026.06 | 26.03 | 24.22 | |
| Qwen-2.5VL-7BInput Modality=BEV Images, Evaluation Protocol=Few-shot Example2026.06 | 22.25 | 25.33 | |
| GPT-4oInput Modality=Raw Images2026.06 | 18.26 | 23.04 | |
| GPT-4oInput Modality=BEV Images2026.06 | 15.49 | 21.91 |