Scene Text Visual Question Answering on ST-VQA
68.96AccuracyQwen 2.5 VL + ViCrop (rel-att)
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen 2.5 VL + ViCrop (rel-att)Base Model=Qwen 2.5 VL, Cropping Strategy=ViCrop (rel-att), Input Resolution=448x4482025.11 | 68.96 | |
| Qwen 2.5 VL + CropVLMBase Model=Qwen 2.5 VL, Cropping Strategy=CropVLM, Input Resolution=2048x20482025.11 | 68.31 | |
| Qwen 2.5 VL + ViCrop (grad-att)Base Model=Qwen 2.5 VL, Cropping Strategy=ViCrop (grad-att), Input Resolution=448x4482025.11 | 68.09 | |
| Qwen 2.5 VL + UV-CoTBase Model=Qwen 2.5 VL, Cropping Strategy=UV-CoT, Input Resolution=336x3362025.11 | 67.91 | |
| Qwen 2.5 VLBase Model=Qwen 2.5 VL, Cropping Strategy=None, Input Resolution=448x4482025.11 | 65.49 | |
| LLaVA 1.5 + UV-CoTBase Model=LLaVA 1.5, Cropping Strategy=UV-CoT, Input Resolution=336x3362025.11 | 59.3 | |
| LLaVA 1.5 + ViCrop (grad-att)Base Model=LLaVA 1.5, Cropping Strategy=ViCrop (grad-att), Input Resolution=336x3362025.11 | 57.06 | |
| LLaVA 1.5 + ViCrop (rel-att)Base Model=LLaVA 1.5, Cropping Strategy=ViCrop (rel-att), Input Resolution=336x3362025.11 | 56.95 | |
| LLaVA 1.5 + CropVLMBase Model=LLaVA 1.5, Cropping Strategy=CropVLM, Input Resolution=2048x20482025.11 | 56.81 | |
| LLaVA 1.5Base Model=LLaVA 1.5, Cropping Strategy=None, Input Resolution=336x3362025.11 | 52.48 |