Visual Question Answering on CrossViewBench 1.0 (test)
86.1Overall AccuracyHumanBase
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| HumanBaseModel Category=Human Performance2026.05 | 86.1 | 87.5 | 80.2 | 86.5 | 93.6 | |
| CrossViewerBackbone=Qwen3-VL-8B-Instruct, Evaluation Protocol=native <REGION> token-replacement2026.05 | 62.7 | 83.2 | 61.1 | 49.1 | 74.4 | |
| Qwen3.5-397BModel Category=Open-source Models2026.05 | 51.7 | 50.1 | 41 | 54.1 | 72.6 | |
| Gemini-3.1-ProModel Category=Proprietary Models2026.05 | 51.5 | 60 | 39 | 50.5 | 56 | |
| Qwen3.5-35BModel Category=Open-source Models2026.05 | 50.1 | 48.3 | 39.4 | 53.3 | 65.6 | |
| GPT-5.2Model Category=Proprietary Models2026.05 | 49.5 | 41.5 | 45.1 | 54.5 | 58.3 | |
| Qwen3-VL-235B-A22B-InstructModel Category=Open-source Models2026.05 | 47.7 | 46.3 | 37.5 | 50.2 | 65.3 | |
| InternVL2.5-38BModel Category=Open-source Models2026.05 | 45.9 | 43.8 | 36.2 | 48.4 | 65.5 | |
| Qwen2.5-VL-72BModel Category=Open-source Models2026.05 | 45.2 | 43.2 | 35.6 | 47.8 | 63.2 | |
| LLaVA-Video-Qwen2-72BModel Category=Open-source Models2026.05 | 44.2 | 42.6 | 34.9 | 46.7 | 60.1 | |
| InternVL2.5-78BModel Category=Open-source Models2026.05 | 43.7 | 42.4 | 34.3 | 46.4 | 56.8 | |
| LLaVA-OneVision-Qwen2-72BModel Category=Open-source Models2026.05 | 43.6 | 42.2 | 34.5 | 45.7 | 61.1 | |
| DeepSeek-VL2Model Category=Open-source Models2026.05 | 42.8 | 41.6 | 33.8 | 44.8 | 59.8 | |
| Qwen3-VL-8BModel Category=Open-source Models2026.05 | 42.7 | 40.1 | 30.7 | 45.3 | 71.1 | |
| Grok-4-FastModel Category=Proprietary Models2026.05 | 42.5 | 38.8 | 33.3 | 45.8 | 59 | |
| InternVL2.5-4BModel Category=Open-source Models2026.05 | 42 | 40.2 | 33.5 | 44.4 | 57.3 | |
| InternVL2.5-2BModel Category=Open-source Models2026.05 | 40 | 38.8 | 31.9 | 42.2 | 52.7 |