Visual Question Answering on Vision-Grounded Study Question-Guided
48.4LLM-judged AccuracyQwen2-VL-72B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2-VL-72BModel Type=open-weight2026.06 | 48.4 | |
| GPT-4o miniModel Type=proprietary/API2026.06 | 46.3 | |
| Gemini 1.5 ProModel Type=proprietary/API2026.06 | 46.1 | |
| Qwen2-VL-7BModel Type=open-weight2026.06 | 45.7 | |
| Claude 3.5 SonnetModel Type=proprietary/API2026.06 | 41.7 | |
| Gemini 1.5 FlashModel Type=proprietary/API2026.06 | 37 | |
| Llama-3.2 11B VisionModel Type=open-weight2026.06 | 23.3 | |
| InternVL2-26BModel Type=open-weight2026.06 | 20.6 | |
| Phi-3.5 VisionModel Type=open-weight2026.06 | 19.1 | |
| LLaVA-OneVision-7BModel Type=open-weight2026.06 | 18.7 | |
| LLaVA-OneVision-0.5BModel Type=open-weight2026.06 | 16.5 |