Visually Grounded Reasoning on V* Bench
95.7Average Accuracyo3-0416
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| o3-0416Category=Private Models2025.07 | 95.7 | — | — | |
| TreeVGR-7BCategory=Open-source Visual Grounded Reasoning Models, Parameters=7B2025.07 | 91.1 | 94 | 87 | |
| DeepScanModel Category=Visually Grounded Reasoning Models, Backbone=Qwen2.5-VL-7B2026.03 | 90.6 | 93 | 86.8 | |
| DeepEyesModel Category=Visually Grounded Reasoning Models, Training Method=RL-based2026.03 | 90 | 92.1 | 86.8 | |
| DeepEyes-7BCategory=Open-source Visual Grounded Reasoning Models, Parameters=7B2025.07 | 90 | 92.1 | 86.8 | |
| Qwen2.5VL-32BModel Category=General Large Vision-Language Models2026.03 | 85.9 | 83.5 | 89.5 | |
| TreeVGRModel Category=Visually Grounded Reasoning Models, Training Method=RL-based, Self-collected=true2026.03 | 85.9 | 86.1 | 85.5 | |
| Qwen2.5VL-72BModel Category=General Large Vision-Language Models2026.03 | 84.8 | 90.8 | 80.9 | |
| Qwen2.5-VL-72BCategory=Open-source General Models, Parameters=72B2025.07 | 84.8 | 90.8 | 80.9 | |
| DyfoModel Category=Visually Grounded Reasoning Models, Self-collected=true2026.03 | 84.3 | 82.6 | 86.8 | |
| Thyme-VLModel Category=Visually Grounded Reasoning Models, Training Method=RL-based2026.03 | 82.2 | 83.5 | 80.3 | |
| ZoomRefineModel Category=Visually Grounded Reasoning Models, Self-collected=true2026.03 | 82.2 | 85.3 | 77.6 | |
| PixelReasonerModel Category=Visually Grounded Reasoning Models, Training Method=RL-based2026.03 | 80.6 | 83.5 | 76.3 | |
| Pixel-Reasoner-7BCategory=Open-source Visual Grounded Reasoning Models, Parameters=7B2025.07 | 80.6 | 83.5 | 76.3 | |
| InternVL3-38BModel Category=General Large Vision-Language Models2026.03 | 77.5 | 77.4 | 77.6 | |
| InternVL3-78BModel Category=General Large Vision-Language Models2026.03 | 76.4 | 75.7 | 77.6 | |
| InternVL3-78BCategory=Open-source General Models, Parameters=78B2025.07 | 76.4 | 75.7 | 77.6 | |
| Qwen2.5-VL-7B + ECRDParameters=7B, Enhancement=ECRD2026.02 | 74.9 | 74.8 | 75 | |
| Qwen2.5VL-7BModel Category=General Large Vision-Language Models2026.03 | 74.3 | 77.4 | 69.7 | |
| Qwen2.5-VL-7BCategory=Open-source General Models, Parameters=7B2025.07 | 74.3 | 77.4 | 69.7 | |
| LLaVA-OV-72BModel Category=General Large Vision-Language Models2026.03 | 73.8 | 80.9 | 63.2 | |
| LLaVA-OneVision-72BCategory=Open-source General Models, Parameters=72B2025.07 | 73.8 | 80.9 | 63.2 | |
| Qwen2.5-VL-7B + supervisorParameters=7B, Enhancement=supervisor2026.02 | 72.8 | 73.9 | 71.1 | |
| InternVL3-8BModel Category=General Large Vision-Language Models2026.03 | 72.3 | 73 | 71.1 | |
| InternVL3-8BCategory=Open-source General Models, Parameters=8B2025.07 | 72.3 | 73 | 71.1 | |
| LLaVA-OneVision-7B + ECRDParameters=7B, Enhancement=ECRD2026.02 | 71.2 | 74.8 | 65.8 | |
| Qwen2.5-VL-7BParameters=7B2026.02 | 71.2 | 73.9 | 67.1 | |
| LLaVA-OV-7BModel Category=General Large Vision-Language Models2026.03 | 70.7 | 73 | 60.5 | |
| LLaVA-OneVision-7BCategory=Open-source General Models, Parameters=7B2025.07 | 70.7 | 73 | 60.5 | |
| LLaVA-OneVision-7BParameters=7B2026.02 | 68.1 | 73 | 60.5 | |
| GPT-4o-1120Model Category=General Large Vision-Language Models2026.03 | 66 | — | — | |
| GPT-4o-1120Category=Private Models2025.07 | 66 | — | — |