Visual Question Answering on HRBench 8K
76.25AccuracyFOCUS
Evaluation Results
| Method | Links | |
|---|---|---|
| FOCUSBase Model=Qwen-2.5-VL2025.06 | 76.25 | |
| DeepEyes-7BThinking Level=Thinking with Images2026.02 | 72.6 | |
| Thyme-7BThinking Level=Thinking with Images2026.02 | 72.4 | |
| ViLAVT-7BThinking Level=Chatting with Images2026.02 | 69.3 | |
| UG-SearchBase Model=LLaVA-OV-7B2025.10 | 68.9 | |
| Qwen-2.5-VL2025.06 | 68.62 | |
| Pixel-Reasoner-7BThinking Level=Thinking with Images2026.02 | 66.9 | |
| ZoomEyeBase Model=LLaVA-OV-7B2025.10 | 66.5 | |
| UG-SearchBase Model=Qwen2.5-VL-7B2025.10 | 66.3 | |
| CropVLMReward=Accuracy, Backbone=Qwen 2.5 VL, Input Resolution=1792x1792, CropVLM Resolution=2048x20482025.11 | 65.63 | |
| Qwen2.5-VL-7BThinking Level=Non-Thinking2026.02 | 65.1 | |
| Qwen2.5-VL-7BThinking Level=Thinking about Images2026.02 | 64.9 | |
| Qwen 2.5 VLInput Resolution=1792x17922025.11 | 63.88 | |
| CropVLMBase Model=Qwen 2.5 VL 3B, Resolution=1024, Cropping Strategy=CropVLM2025.11 | 63.63 | |
| CropVLMReward=LL, Backbone=Qwen 2.5 VL, Input Resolution=1792x1792, CropVLM Resolution=2048x20482025.11 | 63.25 | |
| CropVLMBase Model=Qwen 2.5 VL 3B, Resolution=512, Cropping Strategy=CropVLM2025.11 | 63 | |
| InternVL3-8BThinking Level=Non-Thinking2026.02 | 62 | |
| Qwen 2.5 VL + CropVLMBase Model=Qwen 2.5 VL, Cropping Strategy=CropVLM, Input Resolution=2048x20482025.11 | 60.75 | |
| CropVLMBase Model=Qwen 2.5 VL 3B, Resolution=2048, Cropping Strategy=CropVLM2025.11 | 60.75 | |
| LLaVA-OneVision-7BThinking Level=Non-Thinking2026.02 | 59.8 | |
| ThymeBase Model=Qwen2.5-VL-7B2025.10 | 59 | |
| LLaVA-OV-7BBase Model=LLaVA-OV-7B2025.10 | 58.4 | |
| VILASR-7BThinking Level=Thinking with Images2026.02 | 56.3 | |
| Qwen2.5-VL-7BBase Model=Qwen2.5-VL-7B2025.10 | 53.1 | |
| ViCropBase Model=Qwen2.5-VL-7B2025.10 | 51.9 | |
| TextCoTBase Model=Qwen2.5-VL-7B2025.10 | 50.6 | |
| SpaceR-7BThinking Level=Thinking about Images2026.02 | 49.8 | |
| Qwen 2.5 VL + ViCrop (rel-att)Base Model=Qwen 2.5 VL, Cropping Strategy=ViCrop (rel-att), Input Resolution=448x4482025.11 | 47.38 | |
| Qwen 2.5 VL + UV-CoTBase Model=Qwen 2.5 VL, Cropping Strategy=UV-CoT, Input Resolution=336x3362025.11 | 47.25 | |
| Qwen 2.5 VL + ViCrop (grad-att)Base Model=Qwen 2.5 VL, Cropping Strategy=ViCrop (grad-att), Input Resolution=448x4482025.11 | 46 | |
| MRoPEPositional Encoding Method=MRoPE, DIPE Enhancement=+DIPE2026.03 | 45.13 | |
| Qwen 2.5 VLBase Model=Qwen 2.5 VL, Cropping Strategy=None, Input Resolution=448x4482025.11 | 44.5 | |
| Qwen 2.5 VL 3BBase Model=Qwen 2.5 VL 3B, Resolution=-, Cropping Strategy=None2025.11 | 44.5 | |
| Vanilla RoPEPositional Encoding Method=Vanilla RoPE, DIPE Enhancement=+DIPE2026.03 | 44.12 | |
| MRoPE-IPositional Encoding Method=MRoPE-I, DIPE Enhancement=+DIPE2026.03 | 43.88 | |
| MRoPE-IPositional Encoding Method=MRoPE-I, DIPE Enhancement=Base2026.03 | 43.75 | |
| Vanilla RoPEPositional Encoding Method=Vanilla RoPE, DIPE Enhancement=Base2026.03 | 42.63 | |
| MRoPEPositional Encoding Method=MRoPE, DIPE Enhancement=Base2026.03 | 42.25 | |
| CropVLMBase Model=GPT 4.1 nano, Resolution=2048, Cropping Strategy=CropVLM2025.11 | 40.5 | |
| LLaVA 1.5 + CropVLMBase Model=LLaVA 1.5, Cropping Strategy=CropVLM, Input Resolution=2048x20482025.11 | 39.88 | |
| CropVLMBase Model=LLaVA 1.5 7B, Resolution=2048, Cropping Strategy=CropVLM2025.11 | 39.88 | |
| CropVLMBase Model=GPT 4.1 nano, Resolution=1024, Cropping Strategy=CropVLM2025.11 | 39.5 | |
| CropVLMBase Model=GPT 4.1 nano, Resolution=512, Cropping Strategy=CropVLM2025.11 | 38.88 | |
| CropVLMBase Model=LLaVA 1.5 7B, Resolution=1024, Cropping Strategy=CropVLM2025.11 | 38.38 | |
| LLaVA 1.5 + ViCrop (grad-att)Base Model=LLaVA 1.5, Cropping Strategy=ViCrop (grad-att), Input Resolution=336x3362025.11 | 37.12 | |
| GPT 4.1 nanoBase Model=GPT 4.1 nano, Resolution=-, Cropping Strategy=None2025.11 | 36.88 | |
| LLaVA 1.5 + ViCrop (rel-att)Base Model=LLaVA 1.5, Cropping Strategy=ViCrop (rel-att), Input Resolution=336x3362025.11 | 36.5 | |
| LLaVA 1.5 + UV-CoTBase Model=LLaVA 1.5, Cropping Strategy=UV-CoT, Input Resolution=336x3362025.11 | 35.88 | |
| CropVLMBase Model=LLaVA 1.5 7B, Resolution=512, Cropping Strategy=CropVLM2025.11 | 35.88 | |
| LLaVA 1.5Base Model=LLaVA 1.5, Cropping Strategy=None, Input Resolution=336x3362025.11 | 34.65 | |
| LLaVA 1.5 7BBase Model=LLaVA 1.5 7B, Resolution=-, Cropping Strategy=None2025.11 | 34.65 |