Text-rich Image Question Answering (Extraction) on TRINS-VQA
63.6AccuracyQwen-VL
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen-VLResolution=448^22024.06 | 63.6 | |
| LaRAResolution=336^22024.06 | 62.8 | |
| LLaVAR (finetuned)Resolution=336^2, Fine-tuning=true2024.06 | 61.2 | |
| mPLUG-Owl2Resolution=448^22024.06 | 61 | |
| LLaVAR w/ OCRResolution=336^2, OCR integration=true2024.06 | 58.1 | |
| LLaVARResolution=336^22024.06 | 51.7 | |
| Instruct-BLIPResolution=224^22024.06 | 43.9 | |
| LLaVA 1.5Resolution=336^22024.06 | 38.8 | |
| LLaVAResolution=336^22024.06 | 23.7 | |
| Mini-GPTv2Resolution=224^22024.06 | 15.3 |