Visual Question Answering on TextVQA (Score)
80.2ScorePre-trained
Evaluation Results
| Method | Links | |
|---|---|---|
| Pre-trainedTrainable Params=02026.03 | 80.2 | |
| IADA +LoRATrainable Params=7.7M2026.03 | 77.6 | |
| LoRA-onlyTrainable Params=7.6M2026.03 | 76.2 | |
| AttnRes+LoRATrainable Params=7.7M2026.03 | 76.1 | |
| PinPointTraining Data=GQA, Avg. FLOPs(T)=24.39 (61.1%)2026.03 | 72.93 | |
| LaVerVisual Encoder Model=Qwen-ViT2025.12 | 69.89 | |
| LaVerVisual Encoder Model=SigLIP22025.12 | 69.82 | |
| BaselineVisual Encoder Model=SigLIP22025.12 | 67.87 | |
| LaVerVisual Encoder Model=AIMv22025.12 | 67.04 | |
| LaVerVisual Encoder Model=CLIP2025.12 | 65.58 | |
| BaselineVisual Encoder Model=CLIP2025.12 | 64.39 | |
| BaselineVisual Encoder Model=AIMv22025.12 | 63.7 | |
| BaselineVisual Encoder Model=Qwen-ViT2025.12 | 62.87 | |
| LaVerVisual Encoder Model=DINOv22025.12 | 61.34 | |
| BaselineVisual Encoder Model=DINOv22025.12 | 60.08 | |
| LaVerVisual Encoder Model=MLP + Qwen 2.52025.12 | 54.67 | |
| VanillaTraining Data=-, Avg. FLOPs(T)=39.92 (100.0%)2026.03 | 53.27 | |
| PDropTraining Data=Train-free, Avg. FLOPs(T)=26.00 (65.1%)2026.03 | 52.4 | |
| BaselineVisual Encoder Model=MLP + Qwen 2.52025.12 | 52.29 | |
| FastVTraining Data=Train-free, Avg. FLOPs(T)=28.17 (70.6%)2026.03 | 51.02 |