High-Resolution Visual Perception on HR-Bench 8K
86.88AccuracyGemini-3.1-Pro
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Gemini-3.1-ProParam Size=-, Inference Mode=Closed-Source2026.05 | 86.88 | — | |
| Qwen3.5Param Size=397B, Inference Mode=Single Forward Pass2026.05 | 85.5 | — | |
| Vision-OPD (Ours)Param Size=9B, Inference Mode=Single Forward Pass2026.05 | 85.5 | — | |
| ZwZParam Size=8B, Inference Mode=Single Forward Pass2026.05 | 81.75 | — | |
| Gemini-2.5-ProModel Type=Closed-Source2026.05 | 81.5 | — | |
| SpecReasonBase Model=Thyme2026.03 | 81.02 | 0.51 | |
| Qwen3.5Param Size=9B, Inference Mode=Single Forward Pass2026.05 | 80.63 | — | |
| Qwen3-VL-InstructParam Size=235B, Inference Mode=Single Forward Pass2026.05 | 80.38 | — | |
| Vision-OPD (Ours)Param Size=4B, Inference Mode=Single Forward Pass2026.05 | 80.38 | — | |
| Qwen3.5Param Size=4B, Inference Mode=Single Forward Pass2026.05 | 80.13 | — | |
| GLM-4.6VParam Size=106B, Inference Mode=Single Forward Pass2026.05 | 78.88 | — | |
| SenseNova-MARSParam Size=8B, Inference Mode=Agentic2026.05 | 78.38 | — | |
| GPT-5.2Param Size=-, Inference Mode=Closed-Source2026.05 | 78.38 | — | |
| GPT-5.4Param Size=-, Inference Mode=Closed-Source2026.05 | 77.88 | — | |
| Qwen3-VL-InstructParam Size=8B, Inference Mode=Single Forward Pass2026.05 | 75.25 | — | |
| Kimi-K2.5Param Size=1T, Inference Mode=Single Forward Pass2026.05 | 75.25 | — | |
| MAESTROModel Type=Ours2026.05 | 74.4 | — | |
| GPT-5Model Type=Closed-Source2026.05 | 74.1 | — | |
| DeepEyes-v2Model Type=Think with Images Methods2026.05 | 73.8 | — | |
| DeepEyesV2Param Size=7B, Inference Mode=Agentic2026.05 | 73.75 | — | |
| Gemini-2.5-FlashModel Type=Closed-Source2026.05 | 73.7 | — | |
| SpecEyesBase Model=Thyme, Aggregation Strategy=min2026.03 | 73.31 | 0.95 | |
| GLM-4.6VModel Type=Open-Source & Baselines2026.05 | 73 | — | |
| DeepEyesParam Size=7B, Inference Mode=Agentic2026.05 | 72.63 | — | |
| DeepEyesModel Type=Think with Images Methods2026.05 | 72.6 | — | |
| SpecReasonBase Model=DeepEyes2026.03 | 72.54 | 0.42 | |
| ThymeBase Model=Thyme2026.03 | 72.43 | 1 | |
| SpecEyesBase Model=Thyme, Aggregation Strategy=bottom2026.03 | 72.31 | 0.99 | |
| ThymeParam Size=7B, Inference Mode=Agentic2026.05 | 72 | — | |
| ThymeModel Type=Think with Images Methods2026.05 | 72 | — | |
| SpecEyesBase Model=DeepEyes, Aggregation Strategy=min2026.03 | 71.8 | 1.08 | |
| DeepEyesBase Model=DeepEyes2026.03 | 71.43 | 1 | |
| SpecEyesBase Model=DeepEyes, Aggregation Strategy=bottom2026.03 | 71.18 | 1.04 | |
| SpecEyesBase Model=Thyme, Aggregation Strategy=log2026.03 | 70.84 | 1.06 | |
| MathCoder-VLModel Type=Think with Images Methods2026.05 | 70.6 | — | |
| AdaFocusModel=Qwen2.5VL-3B, Training-free=true2026.02 | 70 | — | |
| SpecEyesBase Model=DeepEyes, Aggregation Strategy=log2026.03 | 69.67 | 1.28 | |
| Qwen3-VL-32BModel Type=Open-Source & Baselines2026.05 | 69.5 | — | |
| MiMo-VL-RLParam Size=7B, Inference Mode=Single Forward Pass2026.05 | 69.38 | — | |
| Untrained ModelModel Type=Open-Source & Baselines2026.05 | 68.5 | — | |
| ZoomEyesModel=Qwen2.5VL-3B, Training-free=true2026.02 | 68.38 | — | |
| Direct AnsweringModel Type=Open-Source & Baselines2026.05 | 68.1 | — | |
| Qwen3-VL-2BMode=draft only2026.03 | 68 | 2.9 | |
| SpecEyesBase Model=Thyme, Aggregation Strategy=mean2026.03 | 68 | 1.21 | |
| Chain-of-FocusModel Type=Think with Images Methods2026.05 | 67.5 | — | |
| SpecEyesBase Model=DeepEyes, Aggregation Strategy=mean2026.03 | 67.38 | 1.77 | |
| VTS-VModel Type=Think with Images Methods2026.05 | 67.3 | — | |
| VisionReasonerModel Type=Think with Images Methods2026.05 | 66.5 | — | |
| VTOOL-R1Model Type=Think with Images Methods2026.05 | 66.4 | — | |
| Pixel ReasonerModel=Qwen2.5VL-3B, Training-free=false2026.02 | 66 | — | |
| PixelReasonerModel Type=Think with Images Methods2026.05 | 65.4 | — | |
| Kimi-K2.5Model Type=Open-Source & Baselines2026.05 | 65.1 | — | |
| MLLMs-KnowModel=Qwen2.5VL-3B, Training-free=true2026.02 | 64.88 | — | |
| MiniCPM-V-4.5Param Size=9B, Inference Mode=Single Forward Pass2026.05 | 61.5 | — | |
| BaselineModel=Qwen2.5VL-3B, Training-free=true2026.02 | 58.88 | — | |
| GPT-4oModel Type=Closed-Source2026.05 | 55 | — | |
| Visual-ARFTModel Type=Think with Images Methods2026.05 | 54 | — | |
| ZoomEyesModel=LLaVA-v1.5-7B, Training-free=true2026.02 | 48.63 | — | |
| AdaFocusModel=LLaVA-v1.5-7B, Training-free=true2026.02 | 42 | — | |
| DC2Model=LLaVA-v1.5-7B, Training-free=true2026.02 | 39.5 | — | |
| MLLMs-KnowModel=LLaVA-v1.5-7B, Training-free=true2026.02 | 37.25 | — | |
| VisCropModel=LLaVA-v1.5-7B, Training-free=true2026.02 | 35.75 | — | |
| BaselineModel=LLaVA-v1.5-7B, Training-free=true2026.02 | 32.13 | — |