Visual Question Answering on Vision-Grounded Study Multi-Signal
58.7LLM-judged AccuracyQwen2-VL-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2-VL-7BModel Type=open-weight2026.06 | 58.7 | |
| Gemini 1.5 ProModel Type=proprietary/API2026.06 | 58.5 | |
| Qwen2-VL-72BModel Type=open-weight2026.06 | 54.8 | |
| Claude 3.5 SonnetModel Type=proprietary/API2026.06 | 54.3 | |
| GPT-4o miniModel Type=proprietary/API2026.06 | 49.4 | |
| Gemini 1.5 FlashModel Type=proprietary/API2026.06 | 43.7 | |
| Llama-3.2 11B VisionModel Type=open-weight2026.06 | 37.6 | |
| InternVL2-26BModel Type=open-weight2026.06 | 36.1 | |
| LLaVA-OneVision-7BModel Type=open-weight2026.06 | 35.9 | |
| LLaVA-OneVision-0.5BModel Type=open-weight2026.06 | 34.6 | |
| Phi-3.5 VisionModel Type=open-weight2026.06 | 29.1 |