Visual Perception and Reasoning on V*
91.1Overall AccuracyTreeVGR-7B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| TreeVGR-7BModel Category=Reasoning with Tool Methods, Backbone=7B2026.06 | 91.1 | 94 | 87 | |
| DeepEyes‡Size=7B, Trajectory Free=true, Maximum visual tokens=16,384, Shuffle options=false2026.03 | 90.1 | 91.3 | 88.2 | |
| DeepEyesModel Category=Reasoning with Tool Methods2026.06 | 90.1 | 91.3 | 88.2 | |
| ActiveScopeModel=Qwen3VL-30B-A3B(MoE)2026.06 | 89.84 | 89.38 | 90.54 | |
| LFPCSize=7B, Trajectory Free=true, Maximum visual tokens=16,3842026.03 | 89.5 | 91.3 | 86.8 | |
| CoF-sftSize=7B, Trajectory Free=false, Maximum visual tokens=16,3842026.03 | 88.5 | 90.4 | 85.5 | |
| ZoomEyeModel=Qwen3VL-30B-A3B(MoE)2026.06 | 88.48 | 90.43 | 85.53 | |
| Mini-o3†Size=7B, Trajectory Free=false, Maximum visual tokens=16,384, Temperature=1.0, Averaging=32 runs2026.03 | 88.2 | — | — | |
| Mini-o3Size=7B, Trajectory Free=false, Maximum visual tokens=16,3842026.03 | 88 | 87.8 | 88.2 | |
| Imagine-OPD-4BModel Category=Our Models, Backbone=4B2026.06 | 88 | 89.6 | 85.5 | |
| Imagine-OPD-8BModel Category=Our Models, Backbone=8B2026.06 | 87.4 | 88.7 | 86.8 | |
| RegularModel=Qwen3VL-30B-A3B(MoE)2026.06 | 86.63 | 86.73 | 86.49 | |
| Qwen3-VL-8BModel Category=Open-Source Models, Backbone=8B2026.06 | 86.4 | 87 | 85.5 | |
| DeepEyesSize=7B, Trajectory Free=true, Maximum visual tokens=16,3842026.03 | 85.9 | 84.3 | 88.2 | |
| Kimi-K2.5 (1T)Model Category=Open-Source Models, Backbone=1T2026.06 | 85.9 | 87 | 84.2 | |
| Pixel ReasonerSize=7B, Trajectory Free=false, Maximum visual tokens=16,3842026.03 | 84.3 | 82.6 | 86.8 | |
| DLRModel Type=Open-Source, Backbone=Qwen3-VL-8B-Thinking2026.04 | 83.8 | 84.3 | 82.9 | |
| RISBackbone=Qwen2.5-VL-7B, Latent Tokens=52026.05 | 83.75 | 84.26 | 83.24 | |
| DeepEyesReasoning Type=Tool-based Visual Reasoning2026.05 | 83.3 | 84.4 | 81.6 | |
| UniVLRReasoning Type=Visual Latent Reasoning2026.05 | 82.7 | 83.5 | 81.6 | |
| LVRModel Type=Open-Source, Backbone=Qwen3-VL-8B-Thinking2026.04 | 82.2 | 82.6 | 81.6 | |
| Qwen3-VL-30B-ThinkingModel Category=Open-Source Models, Backbone=30B2026.06 | 82.2 | 81.7 | 82.9 | |
| ThymeModel Category=Reasoning with Tool Methods2026.06 | 82.2 | 83.5 | 80.3 | |
| MonetBackbone=Qwen2.5-VL-7B, Latent Tokens=52026.05 | 81.9 | 81.94 | 81.86 | |
| RIS+VLPOBackbone=Qwen2.5-VL-7B, Latent Tokens=5, Optimization=Visual-latent Policy Optimization (VLPO)2026.05 | 81.76 | 81.24 | 82.28 | |
| Qwen2.5-VL-32BModel Category=Open-Source Models, Backbone=32B2026.06 | 81.2 | 77.4 | 86.8 | |
| LFPCSize=7B, Trajectory Free=true, Maximum visual tokens=1,0242026.03 | 80.6 | 82.6 | 77.6 | |
| LVRBackbone=Qwen2.5-VL-7B, Latent Tokens=52026.05 | 80.6 | 83.26 | 77.94 | |
| PixelReasonerReasoning Type=Tool-based Visual Reasoning2026.05 | 80.6 | 83.5 | 76.3 | |
| LVRReasoning Type=Visual Latent Reasoning2026.05 | 80.6 | 81.7 | 79.8 | |
| Qwen3-VL-235B-A22BModel Category=Open-Source Models, Backbone=235B-A22B2026.06 | 80.6 | 82.9 | 79.1 | |
| Pixel-ReasonerModel Category=Reasoning with Tool Methods2026.06 | 80.6 | 83.5 | 76.3 | |
| Mini-o3Size=7B, Trajectory Free=false, Maximum visual tokens=1,0242026.03 | 80.1 | 78.3 | 82.9 | |
| PixelReasonerModel Type=Open-Source, Backbone=Qwen3-VL-8B-Thinking2026.04 | 80.1 | 82.6 | 76.3 | |
| SkiLaReasoning Type=Visual Latent Reasoning2026.05 | 80.1 | 79.1 | 81.6 | |
| Qwen3-VL-4BModel Category=Open-Source Models, Backbone=4B2026.06 | 80.1 | 80.9 | 78.9 | |
| Qwen3-VL-8B-ThinkingModel Type=Open-Source, Backbone=Qwen3-VL-8B-Thinking2026.04 | 79.6 | 81.7 | 76.3 | |
| ICoTModel Type=Open-Source, Backbone=Qwen3-VL-8B-Thinking2026.04 | 79.6 | 81.7 | 76.3 | |
| CoVTBackbone=Qwen2.5-VL-7B, Latent Tokens=52026.05 | 79.1 | 81.05 | 77.15 | |
| MonetReasoning Type=Visual Latent Reasoning2026.05 | 79.1 | 81.7 | 75 | |
| Gemini-2.5-ProModel Category=Proprietary Model2026.06 | 79.1 | 86.8 | 68.4 | |
| Qwen2.5-VL-7B+GLSDBackbone=Qwen2.5-VL-7B, Training=Grounded Latent Supervision Dataset (GLSD)2026.05 | 78.25 | 78.39 | 78.11 | |
| ActiveScopeModel=InternVL2.5-26B2026.06 | 78.01 | 80 | 75 | |
| CoVTReasoning Type=Visual Latent Reasoning2026.05 | 78 | 79.1 | 76.3 | |
| Qwen2.5-VL-7BReasoning Type=Textual Reasoning2026.05 | 77.4 | 78.3 | 76.3 | |
| Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B2026.05 | 76.65 | 77.12 | 74.35 | |
| DeepEyesSize=7B, Trajectory Free=true, Maximum visual tokens=1,0242026.03 | 75.9 | 75.7 | 76.3 | |
| Qwen2.5-VL-7B + vanilla SFTReasoning Type=Textual Reasoning, SFT status=vanilla SFT2026.05 | 75.4 | 80 | 68.5 | |
| ZoomEyeModel=InternVL2.5-26B2026.06 | 74.35 | 74.78 | 73.68 | |
| Qwen2.5-VL-7BModel Category=Open-Source Models, Backbone=7B2026.06 | 74.3 | 77.4 | 69.7 | |
| Pixel ReasonerSize=7B, Trajectory Free=false, Maximum visual tokens=1,0242026.03 | 73.9 | 73.1 | 75 | |
| RegularModel=InternVL2.5-26B2026.06 | 73.3 | 73.91 | 72.37 | |
| CoF-sftSize=7B, Trajectory Free=false, Maximum visual tokens=1,0242026.03 | 72.8 | 73 | 72.4 | |
| Gemini-2.5-FlashModel Category=Proprietary Model2026.06 | 72.3 | 77.3 | 64.4 | |
| Gemini-3-FlashModel Category=Proprietary Model2026.06 | 72.3 | 64.4 | 77.3 | |
| InternVL3-8BModel Category=Open-Source Models, Backbone=8B2026.06 | 70.2 | 67.8 | 73.7 | |
| GPT-4oModel Type=Proprietary Model2026.04 | 67.5 | 72.2 | 60.5 | |
| GPT-4oReasoning Type=Textual Reasoning2026.05 | 67.5 | 72.2 | 60.5 | |
| GPT-4oModel Category=Proprietary Model2026.06 | 67.5 | 72.2 | 60.5 | |
| GPT-4oModel Type=Proprietary2026.05 | 65.15 | 69.68 | 58.39 |