High-resolution perception on HR Bench 4K
87.9Overall ScoreGemini-3-Flash
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Gemini-3-FlashModel Category=Proprietary Model2026.04 | 87.9 | — | — | |
| Gemini-2.5-proParams=-2026.02 | 87.87 | 88.75 | 87 | |
| TikArtParams=8B2026.02 | 82.25 | 93.75 | 70.75 | |
| SIEVEBackbone=Qwen3-VL-8B-Instruct2026.03 | 81.5 | 90 | 73 | |
| SIEVEBackbone=Qwen3-VL-4B-Instruct2026.03 | 81.25 | 88 | 74.5 | |
| Qwen3-VL-235B-A22B-InstructParams=235B (22B active)2026.02 | 81.13 | 90.5 | 71.75 | |
| TikArt w/o GRPOParams=8B2026.02 | 81.13 | 92 | 70.25 | |
| Gemini-2.5-flashParams=-2026.02 | 81 | 85 | 77 | |
| GRPOBackbone=Qwen3-VL-8B-Instruct2026.03 | 80.88 | 91.75 | 70 | |
| TikArt w/o ObservationParams=8B2026.02 | 80.63 | 94 | 67.25 | |
| GRPOBackbone=Qwen3-VL-4B-Instruct2026.03 | 80.5 | 88.5 | 72.5 | |
| Qwen3-VL-InstructParams=32B2026.02 | 79.75 | 90.5 | 69 | |
| TikArt w/o Zoom ActionParams=8B2026.02 | 79.5 | 93.75 | 65.25 | |
| Vanilla*Backbone=Qwen3-VL-8B-Instruct2026.03 | 79 | 89 | 69 | |
| DeepEyesV2Model Category=Thinking-with-Images Agent Model2026.04 | 77.9 | — | — | |
| Vanilla*Backbone=Qwen3-VL-4B-Instruct2026.03 | 77.75 | 86 | 69.5 | |
| GPT-5Params=-2026.02 | 77.5 | 78 | 77 | |
| ThymeParams=7B2026.02 | 77 | — | — | |
| Zoom-RefineBackbone=Qwen3-VL-8B-Instruct2026.03 | 77 | 91 | 63 | |
| ThymeModel Category=Thinking-with-Images Agent Model2026.04 | 77 | 91 | 63 | |
| InternVL2.5-8BCVSearch=true2026.05 | 77 | 93 | 61 | |
| InternVL2.5-8BCVSearch=true2026.05 | 77 | 93 | 61 | |
| AutoToolSize=7B2026.05 | 76.9 | 92.5 | 61.3 | |
| Qwen2.5-VL-7BCVSearch=true2026.05 | 76.8 | 91.5 | 61.8 | |
| Qwen2.5-VL-7BCVSearch=true2026.05 | 76.6 | 91.5 | 61.8 | |
| TikArt w/o Segment ActionParams=8B2026.02 | 76.5 | 93.25 | 59.75 | |
| Zoom-RefineBackbone=Qwen3-VL-4B-Instruct2026.03 | 76.5 | 87 | 66 | |
| Qwen2.5-VL-32BSize=32B2026.05 | 76.3 | 87.5 | 65 | |
| InternVL3-38B2026.05 | 76.3 | 83.5 | 69 | |
| LLaVA-OV-7BCVSearch=true2026.05 | 75.6 | 89.5 | 61.8 | |
| LLaVA-OV-7BCVSearch=true2026.05 | 75.6 | 89.5 | 61.8 | |
| ZoomEyeBackbone=Qwen3-VL-4B-Instruct2026.03 | 75.5 | 90 | 61 | |
| SCOLAR-7BBackbone=Qwen2.5-VL-7B2026.05 | 75.5 | 82.5 | 68.5 | |
| MiMo-VL-7B +PCBackbone Model=MiMo-VL-7B, Training Paradigm (+PC)=Yes2026.03 | 75.4 | — | — | |
| MiMo-VL-7BBackbone Model=MiMo-VL-7B, Training Paradigm (+PC)=No2026.03 | 75.2 | — | — | |
| DeepEyesE2E=true, Param Size=7B2025.05 | 75.1 | 91.3 | 59 | |
| DeepEyesModel Category=Thinking-with-Images Agent Model2026.04 | 75.1 | — | — | |
| DeepEyesCVSearch=false2026.05 | 75.1 | 91.3 | 59 | |
| DeepEyesbackbone=Qwen2.5-VL-7B2026.02 | 75.1 | 91.3 | 59 | |
| HyLaR-7BModel Category=Visual-Latent Model2026.04 | 75 | 93.75 | 56.25 | |
| DeepEyesSize=7B2026.05 | 74.9 | 92 | 57.8 | |
| Qwen2.5-VL-32B2026.05 | 74.8 | 89.3 | 60.3 | |
| Qwen3-VL-InstructParams=8B2026.02 | 74.13 | 90.5 | 57.75 | |
| CapImaginebackbone=Qwen2.5-VL-7B2026.02 | 74.1 | 88.5 | 59.8 | |
| CapImaginebackbone=Qwen2.5-VL-7B, rewriting=false2026.02 | 74.1 | 89 | 58.3 | |
| Pixel-ReasonerParams=7B2026.02 | 74 | — | — | |
| ZoomEyeBackbone=Qwen3-VL-8B-Instruct2026.03 | 74 | 88 | 60 | |
| InternVL3-38BSize=38B2026.05 | 74 | 81 | 67 | |
| Qwen2.5-VL*E2E=true, Param Size=32B2025.05 | 73.9 | 89.8 | 58 | |
| Qwen2.5-VL-7BHEE=true2026.07 | 73.9 | 92.3 | 55.5 | |
| Pixel-ReasonerE2E=true, Param Size=7B2025.05 | 72.9 | 86 | 60.3 | |
| Pixel-ReasonerSize=7B2026.05 | 72.9 | 86 | 60.3 | |
| Pixel-ReasonerCVSearch=false2026.05 | 72.9 | 86 | 60.3 | |
| PixelReasonerbackbone=Qwen2.5-VL-7B2026.02 | 72.9 | 86 | 60.3 | |
| LaserModel Category=Visual-Latent Model2026.04 | 72.5 | — | — | |
| SCOLARtraining_stage=SFT only2026.05 | 72.5 | 83.75 | 61.25 | |
| CapImaginebackbone=Qwen2.5-VL-7B, filtering=false2026.02 | 72.5 | 88.3 | 56.8 | |
| InternVL2.5-8BHEE=true2026.07 | 72.4 | 85 | 59.8 | |
| Qwen2.5-VL-7B +PCBackbone Model=Qwen2.5-VL-7B, Training Paradigm (+PC)=Yes2026.03 | 72.3 | — | — | |
| SkiLaModel Category=Visual-Latent Model, Evaluation Pipeline=Reproduced2026.04 | 72.12 | — | — | |
| SkiLaModel Category=Visual-Latent Model, Evaluation Pipeline=Original2026.04 | 72 | — | — | |
| HyLaR-SFTModel Category=Visual-Latent Model2026.04 | 71.5 | 93 | 50 | |
| Qwen2.5-VL-7B2026.05 | 71.5 | 83 | 60 | |
| LLaVA-ov-7BHEE=true2026.07 | 71.5 | 88 | 55 | |
| DeepeyesParams=7B2026.02 | 71.25 | 83.75 | 58.74 | |
| Deepeyes2026.05 | 71.25 | 83.75 | 58.75 | |
| Monet-7BParams=7B2026.02 | 71 | 85.25 | 56.75 | |
| MonetModel Category=Visual-Latent Model, Evaluation Pipeline=Original2026.04 | 71 | 85.25 | 56.75 | |
| Monet2026.05 | 71 | 85.25 | 56.75 | |
| Monetbackbone=Qwen2.5-VL-7B, reported_from_prior_work=true2026.02 | 71 | 85.3 | 56.8 | |
| InternVL3-8Bbackbone=InternVL3-8B2026.02 | 70.8 | 79.3 | 62.3 | |
| LVRbackbone=Qwen2.5-VL-7B2026.02 | 70.8 | 83.8 | 57.8 | |
| Monetbackbone=Qwen2.5-VL-7B, subset=true2026.02 | 70.7 | 88 | 53.3 | |
| Qwen3-VL-ThinkingParams=8B2026.02 | 70.5 | 81 | 60 | |
| InternVL3-8BModel Category=Open-Source Model2026.04 | 70 | 78.8 | 61.3 | |
| InternVL3-8BSize=8B2026.05 | 69.8 | 75.8 | 63.8 | |
| ZoomEyeE2E=false, Param Size=7B2025.05 | 69.6 | 84.3 | 55 | |
| Qwen2.5-VL-7BSize=7B2026.05 | 69.6 | 81.8 | 57.5 | |
| ZoomEyeSize=7B2026.05 | 69.6 | 84.3 | 55 | |
| Qwen2.5-VL*E2E=true, Param Size=7B2025.05 | 68.8 | 85.2 | 52.2 | |
| Qwen2.5-VL-7BBackbone Model=Qwen2.5-VL-7B, Training Paradigm (+PC)=No2026.03 | 68.8 | — | — | |
| Qwen2.5-VL-7BCVSearch=false2026.05 | 68.8 | 85.2 | 52.2 | |
| Qwen2.5-VL-7BCVSearch=false2026.05 | 68.8 | 85.2 | 52.2 | |
| ZoomEyeModel Category=Thinking-with-Images Agent Model2026.04 | 68.75 | 81.25 | 56.25 | |
| Qwen2.5-VL-7BModel Category=Open-Source Model2026.04 | 68 | 80.25 | 55.75 | |
| Qwen2.5VL-7Bbackbone=Qwen2.5VL-7B2026.02 | 68 | 80.3 | 55.8 | |
| DyFoBackbone=Qwen3-VL-8B-Instruct2026.03 | 67.88 | 78.5 | 57.25 | |
| MonetModel Category=Visual-Latent Model, Evaluation Pipeline=Reproduced2026.04 | 67.37 | 78.25 | 56.5 | |
| InternVL2.5-8BHEE=false2026.07 | 66.6 | 76.8 | 56.5 | |
| Qwen2.5-VL-7BHEE=false2026.07 | 66.5 | 81.5 | 51.5 | |
| GPT-4V2026.05 | 66.37 | 71 | 61.75 | |
| Qwen2.5-VL-3BBackbone Model=Qwen2.5-VL-3B, Training Paradigm (+PC)=No2026.03 | 66.3 | — | — | |
| InternVL2.5-8BCVSearch=false2026.05 | 66 | 75.8 | 56.3 | |
| InternVL2.5-8BCVSearch=false2026.05 | 66 | 75.8 | 56.3 | |
| GPT-5-nanoParams=-2026.02 | 65.38 | 69 | 61.75 | |
| Qwen2.5-VL-3B +PCBackbone Model=Qwen2.5-VL-3B, Training Paradigm (+PC)=Yes2026.03 | 65 | — | — | |
| DyFoBackbone=Qwen3-VL-4B-Instruct2026.03 | 65 | 72.75 | 57.25 | |
| GPT-4oSize=-2026.05 | 65 | 66.8 | 63.3 | |
| LLaVA-ov-7BHEE=false2026.07 | 64.6 | 75.3 | 54 | |
| GPT-4oParams=-2026.02 | 63.87 | 68.75 | 59 |