High-Resolution Visual Question Answering on HR-4k
65.13AccuracyQwen 2.5 VL + CropVLM
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen 2.5 VL + CropVLMBase Model=Qwen 2.5 VL, Cropping Strategy=CropVLM, Input Resolution=2048x20482025.11 | 65.13 | |
| Qwen 2.5 VL + ViCrop (rel-att)Base Model=Qwen 2.5 VL, Cropping Strategy=ViCrop (rel-att), Input Resolution=448x4482025.11 | 56.25 | |
| Qwen 2.5 VL + ViCrop (grad-att)Base Model=Qwen 2.5 VL, Cropping Strategy=ViCrop (grad-att), Input Resolution=448x4482025.11 | 54.38 | |
| Qwen 2.5 VL + UV-CoTBase Model=Qwen 2.5 VL, Cropping Strategy=UV-CoT, Input Resolution=336x3362025.11 | 53.62 | |
| Qwen 2.5 VLBase Model=Qwen 2.5 VL, Cropping Strategy=None, Input Resolution=448x4482025.11 | 51.88 | |
| LLaVA 1.5 + ViCrop (rel-att)Base Model=LLaVA 1.5, Cropping Strategy=ViCrop (rel-att), Input Resolution=336x3362025.11 | 44.25 | |
| LLaVA 1.5 + ViCrop (grad-att)Base Model=LLaVA 1.5, Cropping Strategy=ViCrop (grad-att), Input Resolution=336x3362025.11 | 43.62 | |
| LLaVA 1.5 + CropVLMBase Model=LLaVA 1.5, Cropping Strategy=CropVLM, Input Resolution=2048x20482025.11 | 41.38 | |
| LLaVA 1.5 + UV-CoTBase Model=LLaVA 1.5, Cropping Strategy=UV-CoT, Input Resolution=336x3362025.11 | 37.88 | |
| LLaVA 1.5Base Model=LLaVA 1.5, Cropping Strategy=None, Input Resolution=336x3362025.11 | 35.25 |