High-resolution Visual Understanding on HR-Bench 4K
96.5FSPRTWI
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| RTWIModel Variant=Qwen-VL Instruct, Evaluation Setting=Offline2026.02 | 96.5 | 78.8 | — | |
| DeepconfModel Variant=Qwen-VL Instruct, Evaluation Setting=Offline2026.02 | 96 | 76.5 | — | |
| Self-Cer.Model Variant=Qwen-VL Instruct, Evaluation Setting=Offline2026.02 | 96 | 76.5 | — | |
| CISCModel Variant=Qwen-VL Instruct, Evaluation Setting=Offline2026.02 | 95.8 | 77.5 | — | |
| SCModel Variant=Qwen-VL Instruct, Evaluation Setting=Offline2026.02 | 95.3 | 77 | — | |
| BaseModel Variant=Qwen-VL Instruct, Evaluation Setting=Offline2026.02 | 93.8 | 72 | — | |
| ZwZ-Qwen2.5-VL-7B + Q-ZoomRelative inference throughput (Tp)=0.76×2026.04 | 93.8 | 65.3 | 79.5 | |
| ZwZ-Qwen3-VL-4B + Q-ZoomRelative inference throughput (Tp)=0.66×2026.04 | 93 | 69.5 | 81.3 | |
| Qwen2.5-VL-7B + SD-RPNRelative inference throughput (Tp)=0.77×2026.04 | 92.5 | 64.5 | 78.5 | |
| Qwen3-VL-4B + SD-RPNRelative inference throughput (Tp)=0.58×2026.04 | 92.3 | 69.5 | 80.9 | |
| RTWIModel Variant=Qwen-VL Thinking, Evaluation Setting=Offline2026.02 | 91.5 | 65.5 | — | |
| Qwen2.5-VL-7B + DeepEyesEvaluation protocol=Directly cited (†)2026.04 | 91.3 | 59 | 75.1 | |
| Qwen2.5-VL-7B + Q-ZoomRelative inference throughput (Tp)=0.86×2026.04 | 91.3 | 65.8 | 78.5 | |
| Qwen2.5-VL-7B + ThymeRelative inference throughput (Tp)=0.21×2026.04 | 91 | 63 | 77 | |
| Qwen2.5-VL + ThymeBackbone=Qwen2.5-VL, Plugin=Thyme2026.05 | 91 | 63 | 77 | |
| Qwen2.5-VL + V-ABSBackbone=Qwen2.5-VL, Plugin=V-ABS2026.05 | 91 | 67 | 79 | |
| Qwen3-VL + V-ABSBackbone=Qwen3-VL, Plugin=V-ABS2026.05 | 90 | 62 | 76 | |
| Qwen3-VL-4B + Q-ZoomRelative inference throughput (Tp)=0.73×2026.04 | 89.8 | 70.8 | 80.3 | |
| SCModel Variant=Qwen-VL Thinking, Evaluation Setting=Offline2026.02 | 89.5 | 61 | — | |
| DeepconfModel Variant=Qwen-VL Thinking, Evaluation Setting=Offline2026.02 | 89.5 | 62.5 | — | |
| Self-Cer.Model Variant=Qwen-VL Thinking, Evaluation Setting=Offline2026.02 | 89.5 | 61 | — | |
| Qwen2.5-VL-3B + Q-ZoomRelative inference throughput (Tp)=0.67×2026.04 | 88.8 | 54.8 | 71.8 | |
| Qwen2.5-VL-3B + SD-RPNRelative inference throughput (Tp)=0.66×2026.04 | 88 | 58.3 | 73.1 | |
| CISCModel Variant=Qwen-VL Thinking, Evaluation Setting=Offline2026.02 | 87.8 | 59.8 | — | |
| Qwen2.5-VL + ZoomEyeBackbone=Qwen2.5-VL, Plugin=ZoomEye2026.05 | 86.8 | 53.5 | 70.1 | |
| Intern-VL3 + V-ABSBackbone=Intern-VL3, Plugin=V-ABS2026.05 | 86 | 65 | 75.5 | |
| Qwen3-VL-4BRelative inference throughput (Tp)=1.00×2026.04 | 85.8 | 63.8 | 74.8 | |
| ZwZ-Qwen3-VL-4BRelative inference throughput (Tp)=1.00×2026.04 | 85.8 | 66.3 | 76 | |
| Qwen3-VLModel=Qwen3-VL2026.05 | 85 | 55 | 70 | |
| ZwZ-Qwen2.5-VL-7BRelative inference throughput (Tp)=1.00×2026.04 | 84 | 64 | 74 | |
| BaseModel Variant=Qwen-VL Thinking, Evaluation Setting=Offline2026.02 | 83 | 57.3 | — | |
| Qwen2.5-VL-7BModel=Qwen2.5-VL, Parameters=7B2026.05 | 83 | 48 | 65.5 | |
| Qwen2.5-VL-7BRelative inference throughput (Tp)=1.00×2026.04 | 81.8 | 62.8 | 72.5 | |
| GPT-4o + V-ABSBackbone=GPT-4o, Plugin=V-ABS2026.05 | 80.8 | 62 | 71.4 | |
| Qwen2.5-VL-3BRelative inference throughput (Tp)=1.00×2026.04 | 80.5 | 52 | 66.3 | |
| Intern-VL3Model=Intern-VL32026.05 | 76 | 63 | 69.5 | |
| GPT-4oModel=GPT-4o2026.05 | 70 | 48 | 59 | |
| LLaVA-1.5-13B + ViCropRelative inference throughput (Tp)=0.21×2026.04 | 66.3 | 40.3 | 53.3 | |
| LLaVA-1.5-7B + ViCropRelative inference throughput (Tp)=0.15×2026.04 | 60.8 | 34.8 | 47.8 | |
| LLaVA-1.5-13B + Q-ZoomRelative inference throughput (Tp)=0.58×2026.04 | 59.3 | 44.8 | 52 | |
| LLaVA-1.5-7B + SD-RPNRelative inference throughput (Tp)=0.57×2026.04 | 59 | 35.5 | 47.3 | |
| LLaVA-1.5-7B + Q-ZoomRelative inference throughput (Tp)=0.58×2026.04 | 58.8 | 35.8 | 47.1 | |
| LLaVA-1.5-13B + SD-RPNRelative inference throughput (Tp)=0.57×2026.04 | 58.8 | 47 | 52.9 | |
| LLaVA-HR-X-7BModel=LLaVA-HR-X, Parameters=7B2026.05 | 57.8 | 46.3 | 52 | |
| LLaVA-1.5-13B + S2Relative inference throughput (Tp)=0.73×2026.04 | 51.8 | 46.5 | 49.1 | |
| LLaVA-1.5-7B + S2Relative inference throughput (Tp)=0.63×2026.04 | 49.8 | 38.3 | 44 | |
| LLaVA-v1.6-13BModel=LLaVA-v1.6, Parameters=13B2026.05 | 49.8 | 41.3 | 45.5 | |
| LLaVA-1.5-13BRelative inference throughput (Tp)=1.00×2026.04 | 41.5 | 44.5 | 43 | |
| LLaVA-1.5-7BRelative inference throughput (Tp)=1.00×2026.04 | 39.5 | 35.5 | 37.5 | |
| Qwen2.5-VL + DeepEyes v2Backbone=Qwen2.5-VL, Plugin=DeepEyes v22026.05 | — | — | 77.9 | |
| Qwen2.5-VL + Pixel-ReasonerBackbone=Qwen2.5-VL, Plugin=Pixel-Reasoner2026.05 | — | — | 74 | |
| Qwen2.5-VL-7B + DeepEyesv2Evaluation protocol=Directly cited (†)2026.04 | — | — | 77.9 |