Visual Question Answering on Vision-Grounded Study Base Question
74.3LLM-judged AccuracyGemini 1.5 Pro
Evaluation Results
| Method | Links | |
|---|---|---|
| Gemini 1.5 ProModel Type=proprietary/API2026.06 | 74.3 | |
| Qwen2-VL-7BModel Type=open-weight2026.06 | 69.3 | |
| Qwen2-VL-72BModel Type=open-weight2026.06 | 64.8 | |
| Claude 3.5 SonnetModel Type=proprietary/API2026.06 | 64.6 | |
| GPT-4o miniModel Type=proprietary/API2026.06 | 61.1 | |
| Gemini 1.5 FlashModel Type=proprietary/API2026.06 | 58.9 | |
| InternVL2-26BModel Type=open-weight2026.06 | 50.2 | |
| Phi-3.5 VisionModel Type=open-weight2026.06 | 49.8 | |
| LLaVA-OneVision-7BModel Type=open-weight2026.06 | 47.2 | |
| LLaVA-OneVision-0.5BModel Type=open-weight2026.06 | 45.7 | |
| Llama-3.2 11B VisionModel Type=open-weight2026.06 | 39.8 |