element-level commonsense VQA on CommonSketch
93.5PrecisionGPT-4o
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| GPT-4oModel Type=Proprietary VLM2026.03 | 93.5 | 83.2 | 88.1 | 85.5 | |
| Qwen2.5-VLModel Type=Open-source VLM, Prompting Strategy=JSON-formatted predictions2026.03 | 89.8 | 78.2 | 83.6 | 80.2 | |
| mPLUG-Owl3Model Type=Open-source VLM, Prompting Strategy=JSON-formatted predictions2026.03 | 88.3 | 68.6 | 77.2 | 73.9 | |
| SmolVLMModel Type=Open-source VLM, Prompting Strategy=JSON-formatted predictions2026.03 | 80.8 | 34.6 | 48.5 | 52.6 | |
| InternVLModel Type=Open-source VLM, Prompting Strategy=JSON-formatted predictions2026.03 | 80.4 | 89 | 84.5 | 78.9 | |
| MolmoModel Type=Open-source VLM, Prompting Strategy=JSON-formatted predictions2026.03 | 79.8 | 94.9 | 86.7 | 81.2 | |
| LLaVAModel Type=Open-source VLM, Prompting Strategy=binary yes/no prompt2026.03 | 74.9 | 81.9 | 78.2 | 70.6 | |
| BLIPModel Type=Open-source VLM, Prompting Strategy=binary yes/no prompt2026.03 | 73.1 | 76 | 74.6 | 66.6 | |
| PaliGemma2Model Type=Open-source VLM, Prompting Strategy=JSON-formatted predictions2026.03 | 67.3 | 80.9 | 73.5 | 62.4 |