Visual Question Answering on Vision-Grounded Study Vision-Grounded variant
42LLM-judged AccuracyQwen2-VL-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2-VL-7BModel Type=open-weight2026.06 | 42 | |
| Qwen2-VL-72BModel Type=open-weight2026.06 | 38.4 | |
| Gemini 1.5 ProModel Type=proprietary/API2026.06 | 35.9 | |
| GPT-4o miniModel Type=proprietary/API2026.06 | 33.3 | |
| Claude 3.5 SonnetModel Type=proprietary/API2026.06 | 31.6 | |
| Gemini 1.5 FlashModel Type=proprietary/API2026.06 | 27.3 | |
| InternVL2-26BModel Type=open-weight2026.06 | 15.9 | |
| Llama-3.2 11B VisionModel Type=open-weight2026.06 | 15.2 | |
| LLaVA-OneVision-7BModel Type=open-weight2026.06 | 13.2 | |
| LLaVA-OneVision-0.5BModel Type=open-weight2026.06 | 12.6 | |
| Phi-3.5 VisionModel Type=open-weight2026.06 | 10 |