Multimodal Intent Disambiguation on VAGUE
73.05Direct AccuracyGemini-3.0-Flash
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Gemini-3.0-FlashModel Category=Reasoning VLMs2025.07 | 73.05 | 0.78 | 13.05 | 4.35 | |
| Claude-3.5-SonnetModel Category=Standard VLMs2025.07 | 62.92 | 3.4 | 10.9 | 4.38 | |
| GPT-4oModel Category=Standard VLMs2025.07 | 61.6 | 1.6 | 11.48 | 5.83 | |
| OpenAI o1-fullModel Category=Reasoning VLMs2025.07 | 59.82 | 2.75 | 10.98 | 1.77 | |
| GPT-5.2Model Category=Reasoning VLMs, Reasoning Effort=High2025.07 | 59.21 | 0.78 | 11.9 | 2.81 | |
| Gemini-2.5-ProModel Category=Standard VLMs2025.07 | 53.25 | 5.07 | 6.55 | 14.37 | |
| LLaVA-OneVision-7BModel Category=Open-Source VLMs2025.07 | 49.73 | 3.58 | 5.43 | 4.92 | |
| Qwen2-VL-7BModel Category=Open-Source VLMs2025.07 | 38.52 | 0.03 | 4.5 | 5.9 | |
| OpenAI o3-miniModel Category=Reasoning VLMs2025.07 | 38.15 | 2.06 | 3.72 | 5.49 |