Visual perception and grounding on InfoVQA
88.3AccuracyLlava-OV
Evaluation Results
| Method | Links | |
|---|---|---|
| Llava-OVSize=7B2026.05 | 88.3 | |
| DeepEyesSize=7B2026.05 | 87.7 | |
| MoCASize=7B2026.05 | 87 | |
| Pixel ReasonerSize=7B2026.05 | 86.4 | |
| Qwen2.5-VL-InstructSize=72B2026.05 | 84.3 | |
| GPT-4o-miniSize=-2026.05 | 83.3 | |
| GPT-4oSize=-2026.05 | 80.7 | |
| Qwen2.5-VL-InstructSize=7B2026.05 | 80.7 | |
| VL-RethinkerSize=7B2026.05 | 79.5 | |
| R1-VLSize=7B2026.05 | 78 | |
| mPLUG-Owl3Size=7B2026.05 | 76.3 | |
| DocopilotSize=8B2026.05 | 75 | |
| Claude-3.5Size=-2026.05 | 74.3 |