Visual Question Answering on Vision-Grounded Study Subquestion-Guided variant
56.8LLM-judged AccuracyQwen2-VL-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2-VL-7BModel Type=open-weight2026.06 | 56.8 | |
| Qwen2-VL-72BModel Type=open-weight2026.06 | 54.6 | |
| Gemini 1.5 ProModel Type=proprietary/API2026.06 | 53.3 | |
| GPT-4o miniModel Type=proprietary/API2026.06 | 49.4 | |
| Claude 3.5 SonnetModel Type=proprietary/API2026.06 | 47.8 | |
| Gemini 1.5 FlashModel Type=proprietary/API2026.06 | 41.8 | |
| Llama-3.2 11B VisionModel Type=open-weight2026.06 | 31.3 | |
| LLaVA-OneVision-7BModel Type=open-weight2026.06 | 29.8 | |
| LLaVA-OneVision-0.5BModel Type=open-weight2026.06 | 29.8 | |
| InternVL2-26BModel Type=open-weight2026.06 | 28.9 | |
| Phi-3.5 VisionModel Type=open-weight2026.06 | 22.6 |