Visual Question Answering on HRBench-4K
0.7925AccuracyFOCUS
Evaluation Results
| Method | Links | |
|---|---|---|
| FOCUSBase Model=Qwen-2.5-VL2025.06 | 0.7925 | |
| Thyme-7BThinking Level=Thinking with Images2026.02 | 0.783 | |
| ViLAVT-7BThinking Level=Chatting with Images2026.02 | 0.755 | |
| DeepEyes-7BThinking Level=Thinking with Images2026.02 | 0.751 | |
| UG-SearchBase Model=Qwen2.5-VL-7B2025.10 | 0.749 | |
| Baseline (No Recompute)Model=Qwen3-VL-8B, Recomputation budget (k)=02026.03 | 0.7488 | |
| Qwen3VL-32B-Instruct2026.05 | 0.745 | |
| CLAIMDIFF-RL actor-onlyReward=actor-only, Initialization=Qwen2026.05 | 0.745 | |
| Pixel-Reasoner-7BThinking Level=Thinking with Images2026.02 | 0.729 | |
| No RecomputeModel=Qwen3-VL-8B, Recomputation budget (k)=22026.03 | 0.7275 | |
| InfoFlow KVModel=Qwen3-VL-8B, Recomputation budget (k)=22026.03 | 0.7263 | |
| EPICModel=Qwen3-VL-8B, Recomputation budget (k)=22026.03 | 0.725 | |
| CacheBlendModel=Qwen3-VL-8B, Recomputation budget (k)=22026.03 | 0.7238 | |
| Qwen-2.5-VL2025.06 | 0.7162 | |
| HART-7BResolution=4, 023 × 3, 5032026.02 | 0.711 | |
| InternVL3-8BThinking Level=Non-Thinking2026.02 | 0.708 | |
| InternVL3-7BResolution=4, 023 × 3, 5032026.02 | 0.708 | |
| Qwen2.5-VL-7BResolution=4, 023 × 3, 5032026.02 | 0.701 | |
| UG-SearchBase Model=LLaVA-OV-7B2025.10 | 0.701 | |
| Qwen2.5-VL-7BThinking Level=Thinking about Images2026.02 | 0.698 | |
| InfoFlow KVModel=Qwen3-VL-8B, Recomputation budget (k)=42026.03 | 0.6913 | |
| No RecomputeModel=Qwen3-VL-8B, Recomputation budget (k)=42026.03 | 0.6888 | |
| CacheBlendModel=Qwen3-VL-8B, Recomputation budget (k)=42026.03 | 0.6863 | |
| ZoomEyeBase Model=LLaVA-OV-7B2025.10 | 0.684 | |
| Qwen2.5-VL-7BThinking Level=Non-Thinking2026.02 | 0.678 | |
| Qwen 2.5 VLInput Resolution=1792x17922025.11 | 0.6775 | |
| EPICModel=Qwen3-VL-8B, Recomputation budget (k)=42026.03 | 0.6725 | |
| ThymeBase Model=Qwen2.5-VL-7B2025.10 | 0.666 | |
| CLAIMDIFF-RL relativeReward=relative, Initialization=SFT2026.05 | 0.665 | |
| CropVLMReward=LL, Backbone=Qwen 2.5 VL, Input Resolution=1792x1792, CropVLM Resolution=2048x20482025.11 | 0.6638 | |
| CropVLMReward=Accuracy, Backbone=Qwen 2.5 VL, Input Resolution=1792x1792, CropVLM Resolution=2048x20482025.11 | 0.6575 | |
| CropVLMBase Model=Qwen 2.5 VL 3B, Resolution=2048, Cropping Strategy=CropVLM2025.11 | 0.6513 | |
| LLaVA-OV-7BBase Model=LLaVA-OV-7B2025.10 | 0.649 | |
| CropVLMBase Model=Qwen 2.5 VL 3B, Resolution=512, Cropping Strategy=CropVLM2025.11 | 0.6475 | |
| CropVLMBase Model=Qwen 2.5 VL 3B, Resolution=1024, Cropping Strategy=CropVLM2025.11 | 0.6475 | |
| LLaVA-OneVision-7BThinking Level=Non-Thinking2026.02 | 0.643 | |
| LLaVA-OneVision-7BResolution=4, 023 × 3, 5032026.02 | 0.643 | |
| HOLISTIC-RLReference Setting=w/ ref, Initialization=SFT2026.05 | 0.64 | |
| ViCropBase Model=Qwen2.5-VL-7B2025.10 | 0.633 | |
| CLAIMDIFF-RL actor-onlyReward=actor-only, Initialization=SFT2026.05 | 0.62 | |
| TextCoTBase Model=Qwen2.5-VL-7B2025.10 | 0.606 | |
| VILASR-7BThinking Level=Thinking with Images2026.02 | 0.605 | |
| Qwen2.5-VL-7BBase Model=Qwen2.5-VL-7B2025.10 | 0.601 | |
| SpaceR-7BThinking Level=Thinking about Images2026.02 | 0.581 | |
| HOLISTIC-RLReference Setting=w/o ref, Initialization=SFT2026.05 | 0.575 | |
| MRoPEPositional Encoding Method=MRoPE, DIPE Enhancement=+DIPE2026.03 | 0.545 | |
| SFT2026.05 | 0.535 | |
| MRoPE-IPositional Encoding Method=MRoPE-I, DIPE Enhancement=+DIPE2026.03 | 0.5212 | |
| Qwen 2.5 VL 3BBase Model=Qwen 2.5 VL 3B, Resolution=-, Cropping Strategy=None2025.11 | 0.5188 | |
| MRoPE-IPositional Encoding Method=MRoPE-I, DIPE Enhancement=Base2026.03 | 0.5162 | |
| Vanilla RoPEPositional Encoding Method=Vanilla RoPE, DIPE Enhancement=+DIPE2026.03 | 0.51 | |
| Vanilla RoPEPositional Encoding Method=Vanilla RoPE, DIPE Enhancement=Base2026.03 | 0.5075 | |
| MRoPEPositional Encoding Method=MRoPE, DIPE Enhancement=Base2026.03 | 0.4763 | |
| CropVLMBase Model=LLaVA 1.5 7B, Resolution=1024, Cropping Strategy=CropVLM2025.11 | 0.4388 | |
| CropVLMBase Model=GPT 4.1 nano, Resolution=2048, Cropping Strategy=CropVLM2025.11 | 0.4313 | |
| CropVLMBase Model=LLaVA 1.5 7B, Resolution=2048, Cropping Strategy=CropVLM2025.11 | 0.4138 | |
| CropVLMBase Model=GPT 4.1 nano, Resolution=1024, Cropping Strategy=CropVLM2025.11 | 0.4138 | |
| CropVLMBase Model=LLaVA 1.5 7B, Resolution=512, Cropping Strategy=CropVLM2025.11 | 0.3988 | |
| GPT 4.1 nanoBase Model=GPT 4.1 nano, Resolution=-, Cropping Strategy=None2025.11 | 0.3875 | |
| CropVLMBase Model=GPT 4.1 nano, Resolution=512, Cropping Strategy=CropVLM2025.11 | 0.3863 | |
| LLaVA 1.5 7BBase Model=LLaVA 1.5 7B, Resolution=-, Cropping Strategy=None2025.11 | 0.3525 |