High-resolution Visual Understanding on HR-Bench 8K
95FSPRTWI
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| RTWIModel Variant=Qwen-VL Instruct, Evaluation Setting=Offline2026.02 | 95 | 76 | — | — | |
| SCModel Variant=Qwen-VL Instruct, Evaluation Setting=Offline2026.02 | 93 | 74.8 | — | — | |
| CISCModel Variant=Qwen-VL Instruct, Evaluation Setting=Offline2026.02 | 93 | 75.8 | — | — | |
| DeepconfModel Variant=Qwen-VL Instruct, Evaluation Setting=Offline2026.02 | 93 | 75.8 | — | — | |
| Self-Cer.Model Variant=Qwen-VL Instruct, Evaluation Setting=Offline2026.02 | 92.5 | 75.5 | — | — | |
| ZwZ-Qwen2.5-VL-7B + Q-ZoomRelative inference throughput (Tp)=0.76×2026.04 | 92.5 | 66.3 | 79.4 | — | |
| Zoom-RefineBackbone=Qwen3-VL-8B-Instruct2026.03 | 92 | 60 | 76 | — | |
| ZwZ-Qwen3-VL-4B + Q-ZoomRelative inference throughput (Tp)=0.66×2026.04 | 91.3 | 69.8 | 80.5 | — | |
| BaseModel Variant=Qwen-VL Instruct, Evaluation Setting=Offline2026.02 | 90.8 | 71.3 | — | — | |
| Qwen2.5-VL-7B + Q-ZoomRelative inference throughput (Tp)=0.86×2026.04 | 90.8 | 63.8 | 77.3 | — | |
| TikArt w/o Segment ActionParams=8B2026.02 | 89.25 | 61.5 | 75.38 | — | |
| RTWIModel Variant=Qwen-VL Thinking, Evaluation Setting=Offline2026.02 | 89 | 63 | — | — | |
| CISCModel Variant=Qwen-VL Thinking, Evaluation Setting=Offline2026.02 | 88.5 | 56 | — | — | |
| ZoomEyeE2E=false, Param Size=7B2025.05 | 88.5 | 50 | 69.3 | — | |
| HyLaR-7BModel Category=Visual-Latent Model2026.04 | 88.25 | 52.75 | 70.5 | — | |
| Self-Cer.Model Variant=Qwen-VL Thinking, Evaluation Setting=Offline2026.02 | 88 | 59 | — | — | |
| ZoomEyeBackbone=Qwen3-VL-4B-Instruct2026.03 | 88 | 60 | 74 | — | |
| ZoomEyeBackbone=Qwen3-VL-8B-Instruct2026.03 | 88 | 53 | 70.5 | — | |
| Qwen3-VL-4B + SD-RPNRelative inference throughput (Tp)=0.58×2026.04 | 87.5 | 64.8 | 76.1 | — | |
| Qwen3-VL-4B + Q-ZoomRelative inference throughput (Tp)=0.73×2026.04 | 87.5 | 64.8 | 76.1 | — | |
| Qwen2.5-VL-3B + Q-ZoomRelative inference throughput (Tp)=0.67×2026.04 | 87.3 | 55.5 | 71.4 | — | |
| HyLaR-SFTModel Category=Visual-Latent Model2026.04 | 87.01 | 47.11 | 67.19 | — | |
| DeepconfModel Variant=Qwen-VL Thinking, Evaluation Setting=Offline2026.02 | 87 | 61.8 | — | — | |
| Zoom-RefineBackbone=Qwen3-VL-4B-Instruct2026.03 | 87 | 63 | 75 | — | |
| DeepEyesE2E=true, Param Size=7B2025.05 | 86.8 | 58.5 | 72.6 | — | |
| GRPOBackbone=Qwen3-VL-8B-Instruct2026.03 | 86.75 | 68.25 | 77.5 | — | |
| Qwen2.5-VL-7B + DeepEyesEvaluation protocol=Directly cited (†)2026.04 | 86.6 | 58.5 | 72.6 | — | |
| SCModel Variant=Qwen-VL Thinking, Evaluation Setting=Offline2026.02 | 86.5 | 61 | — | — | |
| Qwen2.5-VL-7B + ThymeRelative inference throughput (Tp)=0.21×2026.04 | 86.5 | 57.5 | 72 | — | |
| Qwen2.5-VL-7B + SD-RPNRelative inference throughput (Tp)=0.77×2026.04 | 86.5 | 60.5 | 73.5 | — | |
| ThymeModel Category=Thinking-with-Images Agent Model2026.04 | 86.5 | 57.5 | 72 | — | |
| TikArt w/o ObservationParams=8B2026.02 | 85.25 | 64.75 | 75 | — | |
| TikArt w/o Zoom ActionParams=8B2026.02 | 85.25 | 64.25 | 74.75 | — | |
| SIEVEBackbone=Qwen3-VL-4B-Instruct2026.03 | 85 | 67.25 | 76.13 | — | |
| Qwen3-VL-InstructParams=32B2026.02 | 84.75 | 68 | 76.38 | — | |
| Gemini-2.5-proParams=-2026.02 | 84.5 | 82.75 | 83.62 | — | |
| TikArtParams=8B2026.02 | 84.5 | 68.25 | 76.38 | — | |
| Qwen2.5-VL*E2E=true, Param Size=32B2025.05 | 84.5 | 56.3 | 70.4 | — | |
| GRPOBackbone=Qwen3-VL-4B-Instruct2026.03 | 83 | 67.75 | 75.38 | — | |
| SIEVEBackbone=Qwen3-VL-8B-Instruct2026.03 | 83 | 73.5 | 78.25 | — | |
| Qwen3-VL-235B-A22B-InstructParams=235B (22B active)2026.02 | 82.25 | 70 | 76.13 | — | |
| Qwen3-VL-InstructParams=8B2026.02 | 81.25 | 55.75 | 68.5 | — | |
| Vanilla*Backbone=Qwen3-VL-8B-Instruct2026.03 | 81 | 67.5 | 74.25 | — | |
| Vanilla*Backbone=Qwen3-VL-4B-Instruct2026.03 | 80.75 | 64 | 72.38 | — | |
| Pixel-ReasonerE2E=true, Param Size=7B2025.05 | 80 | 54.3 | 66.9 | — | |
| Monet-7BParams=7B2026.02 | 79.75 | 56.25 | 68 | — | |
| MonetModel Category=Visual-Latent Model, Evaluation Pipeline=Original2026.04 | 79.75 | 56.25 | 68 | — | |
| Qwen2.5-VL-3B + SD-RPNRelative inference throughput (Tp)=0.66×2026.04 | 79.5 | 51 | 65.3 | — | |
| Gemini-2.5-flashParams=-2026.02 | 79 | 73.5 | 76.25 | — | |
| Qwen2.5-VL*E2E=true, Param Size=7B2025.05 | 78.8 | 51.8 | 65.3 | — | |
| InternVL3-8BModel Category=Open-Source Model2026.04 | 78.8 | 59.8 | 69.3 | — | |
| BaseModel Variant=Qwen-VL Thinking, Evaluation Setting=Offline2026.02 | 78.3 | 56 | — | — | |
| DeepeyesParams=7B2026.02 | 77 | 53.25 | 65.13 | — | |
| ZwZ-Qwen3-VL-4BRelative inference throughput (Tp)=1.00×2026.04 | 77 | 65.8 | 71.4 | — | |
| TikArt w/o GRPOParams=8B2026.02 | 76.5 | 70.25 | 73.38 | — | |
| ZoomEyeModel Category=Thinking-with-Images Agent Model2026.04 | 75 | 54 | 64.5 | — | |
| MonetModel Category=Visual-Latent Model, Evaluation Pipeline=Reproduced2026.04 | 74.25 | 54.49 | 64.37 | — | |
| ZwZ-Qwen2.5-VL-7BRelative inference throughput (Tp)=1.00×2026.04 | 73.8 | 57 | 65.4 | — | |
| Qwen3-VL-ThinkingParams=8B2026.02 | 73.75 | 60.5 | 67.13 | — | |
| Qwen2.5-VL-7BModel Category=Open-Source Model2026.04 | 73.75 | 53.75 | 63.75 | — | |
| Qwen3-VL-4BRelative inference throughput (Tp)=1.00×2026.04 | 73.5 | 64 | 68.8 | — | |
| Qwen2.5-VL-7BRelative inference throughput (Tp)=1.00×2026.04 | 72.8 | 54.5 | 63.6 | — | |
| DyFoBackbone=Qwen3-VL-8B-Instruct2026.03 | 71.5 | 54.5 | 63 | — | |
| GPT-5Params=-2026.02 | 70.5 | 72.75 | 71.62 | — | |
| Qwen2.5-VL-3BRelative inference throughput (Tp)=1.00×2026.04 | 70 | 47.8 | 58.9 | — | |
| DyFoBackbone=Qwen3-VL-4B-Instruct2026.03 | 69.75 | 53.5 | 61.62 | — | |
| LLaVA-OneVision-7BModel Category=Open-Source Model2026.04 | 67.5 | 49 | 58.25 | — | |
| LLaVA-OneVisionE2E=true, Param Size=7B2025.05 | 67.3 | 52.3 | 59.8 | — | |
| GPT-5-nanoParams=-2026.02 | 65.75 | 61.5 | 63.63 | — | |
| GPT-4oE2E=true, Param Size=-2025.05 | 62 | 49 | 55.5 | — | |
| GPT-4oModel Category=Proprietary Model2026.04 | 62 | 49 | 55.5 | — | |
| GPT-4oParams=-2026.02 | 60.75 | 60 | 60.38 | — | |
| LLaVA-OneVision-1.5Params=8B2026.02 | 59.5 | 43.75 | 51.62 | — | |
| LLaVA-1.5-13B + SD-RPNRelative inference throughput (Tp)=0.57×2026.04 | 51.8 | 39.8 | 45.8 | — | |
| LLaVA-1.5-7B + SD-RPNRelative inference throughput (Tp)=0.57×2026.04 | 48.8 | 34.5 | 41.6 | — | |
| LLaVA-1.5-7B + Q-ZoomRelative inference throughput (Tp)=0.58×2026.04 | 48.5 | 34.5 | 41.4 | — | |
| LLaVA-1.5-13B + ViCropRelative inference throughput (Tp)=0.21×2026.04 | 44.5 | 37 | 40.6 | — | |
| LLaVA-1.5-13B + S2Relative inference throughput (Tp)=0.73×2026.04 | 41.3 | 44.5 | 42.9 | — | |
| LLaVA-1.5-7B + S2Relative inference throughput (Tp)=0.63×2026.04 | 40.5 | 36.5 | 38.5 | — | |
| LLaVA-1.5-7B + ViCropRelative inference throughput (Tp)=0.15×2026.04 | 38.5 | 33.8 | 36.1 | — | |
| LLaVA-1.5-13B + Q-ZoomRelative inference throughput (Tp)=0.58×2026.04 | 36.8 | 51 | 43.9 | — | |
| LLaVA-1.5-13BRelative inference throughput (Tp)=1.00×2026.04 | 36.6 | 38.5 | 36.6 | — | |
| LLaVA-1.5-7BRelative inference throughput (Tp)=1.00×2026.04 | 32.5 | 33.8 | 33.8 | — | |
| DeepEyesModel Category=Thinking-with-Images Agent Model2026.04 | — | — | 72.6 | — | |
| DeepEyesBase Model=Qwen2.5-VL-7B2026.05 | — | — | — | 72.6 | |
| DeepEyes-v2Base Model=Qwen2.5-VL-7B2026.05 | — | — | — | 73.8 | |
| DeepEyesV2Model Category=Thinking-with-Images Agent Model2026.04 | — | — | 73.8 | — | |
| Gemini-3-FlashModel Category=Proprietary Model2026.04 | — | — | 85 | — | |
| Mini-o3Base Model=Qwen2.5-VL-7B2026.05 | — | — | — | 73.3 | |
| Pixel-ReasonerParams=7B2026.02 | — | — | 66.9 | — | |
| PixelReasonerBase Model=Qwen2.5-VL-7B2026.05 | — | — | — | 66.9 | |
| PyVision-RLBase Model=Qwen2.5-VL-7B2026.05 | — | — | — | 74.3 | |
| Qwen2.5-VL-7B + DeepEyesv2Evaluation protocol=Directly cited (†)2026.04 | — | — | 73.8 | — | |
| Qwen2.5-VL-7B-InstructBase Model=Qwen2.5-VL-7B2026.05 | — | — | — | 67.9 | |
| Qwen3-VL-8B-Thinking (Agent)Base Model=Qwen3-VL-8B-Thinking2026.05 | — | — | — | 66.1 | |
| SFT + AXPOBase Model=Qwen3-VL-8B-Thinking2026.05 | — | — | — | 77 | |
| SkiLaModel Category=Visual-Latent Model, Evaluation Pipeline=Reproduced2026.04 | — | — | 66.5 | — | |
| SkiLaModel Category=Visual-Latent Model, Evaluation Pipeline=Original2026.04 | — | — | 66.5 | — | |
| ThymeParams=7B2026.02 | — | — | 72 | — | |
| ThymeBase Model=Qwen2.5-VL-7B2026.05 | — | — | — | 72 |