Visual Question Answering on OCR-VQA (test)
77.8AccuracyPaLI-3 (Specialist)
Evaluation Results
| Method | Links | |
|---|---|---|
| PaLI-3 (Specialist)OCR pipeline input=true, Model Category=Specialist, Reference Citation=[8]2023.12 | 77.8 | |
| SOTA [8]2023.12 | 77.8 | |
| PaLI-3OCR Usage=SOTA Overall2024.02 | 77.8 | |
| PaLI-3OCR Usage=With OCR2024.02 | 77.8 | |
| PaLI-XOCR pipeline input=true2023.05 | 77.3 | |
| PaLI-XBase model=PaLI-X2023.12 | 77.3 | |
| PaLI-3 (Specialist)OCR pipeline input=false, Model Category=Specialist, Reference Citation=[8]2023.12 | 76.7 | |
| PaLI-3OCR Usage=Without OCR2024.02 | 76.7 | |
| ScreenAIOCR Usage=With OCR2024.02 | 76.2 | |
| SMoLA-PaLI-X_FT (Specialist)Base model=PaLI-X, Per-task LoRA tuning=true, LoRA rank=42023.12 | 75.7 | |
| PaLI-XOCR pipeline input=false2023.05 | 75 | |
| ScreenAIOCR Usage=Without OCR2024.02 | 75 | |
| SMoLA-PaLI-XFTOCR pipeline input=true, Model Category=Generalist, Backbone=PaLI-X2023.12 | 74.9 | |
| SMoLA-PaLI-3FTOCR pipeline input=true, Model Category=Generalist, Backbone=PaLI-32023.12 | 73.9 | |
| CogVLM-ChatLLM=Vicuna-7B, Trained during SFT stage=true2023.11 | 73.8 | |
| SMoLA-PaLI-3FTOCR pipeline input=false, Model Category=Generalist, Backbone=PaLI-32023.12 | 72.8 | |
| SMoLA-PaLI-XFTOCR pipeline input=false, Model Category=Generalist, Backbone=PaLI-X2023.12 | 71.6 | |
| DocFormerv2_largepre-train data=64M, number of parameters=750M2023.06 | 71.5 | |
| Method [27]OCR pipeline input=false2023.05 | 71.3 | |
| Pix2Struct LargePretraining=Screenshot parsing, Pixel only=true2022.10 | 71.3 | |
| Qwen-VL-ChatLLM=Qwen-7B, Trained during SFT stage=true2023.11 | 70.5 | |
| Qwen-VL-ChatResolution=448x448, In-house data=true2024.03 | 70.5 | |
| GIT2+pre-train data=12.9B, number of parameters=5.1B, extra VQA data=true2023.06 | 70.3 | |
| DocFormerv2_basepre-train data=64M, number of parameters=232M2023.06 | 70.3 | |
| GIT2Pretraining=Image captioning, Pixel only=true2022.10 | 70.3 | |
| GIT22022.05 | 70.3 | |
| GIT2Model Size=5.1B2022.05 | 70.3 | |
| Pix2Struct BasePretraining=Screenshot parsing, Pixel only=true2022.10 | 69.4 | |
| GITpre-train data=800M, number of parameters=681M2023.06 | 68.1 | |
| GIT2022.05 | 68.1 | |
| GIT2022.05 | 68.1 | |
| GITModel Size=0.7B2022.05 | 68.1 | |
| GIT2022.05 | 68.1 | |
| LaTr-BaseOCR System=Rosetta OCR, Pre-training Dataset=IDL, Setting=Constrained2021.12 | 67.9 | |
| LaTr_basepre-train data=64M, number of parameters=311M2023.06 | 67.9 | |
| Prior SOTA2022.05 | 67.9 | |
| LaTr2022.05 | 67.9 | |
| LaTr2022.05 | 67.9 | |
| LaTr2022.05 | 67.9 | |
| SPHINX-2kLLM=LLaMA2 13B, Trained during SFT stage=true2023.11 | 67.8 | |
| Sphinx-2KResolution=768x768, In-house data=false2024.03 | 67.8 | |
| Method [53]OCR pipeline input=true2023.05 | 67.5 | |
| LATtype=State of the art w/ pipelines2022.10 | 67.5 | |
| DonutPretraining=OCR, Pixel only=true2022.10 | 66 | |
| InfiMM-HDResolution=dynamic, In-house data=false2024.03 | 66 | |
| LaAP2023.06 | 64.1 | |
| LaAP-Net2022.05 | 64.1 | |
| LaAP-Net2022.05 | 64.1 | |
| LaAP-Net2022.05 | 64.1 | |
| M4COCR System=Rosetta OCR, Setting=Constrained2021.12 | 63.9 | |
| M4Cnumber of parameters=200M2023.06 | 63.9 | |
| M4C2022.05 | 63.9 | |
| M4C2022.05 | 63.9 | |
| M4C2022.05 | 63.9 | |
| GIT_largepre-train data=20M, number of parameters=347M2023.06 | 62.9 | |
| GIT_LModel Size=0.3B2022.05 | 62.9 | |
| SAIL-VLModel Size=8B2025.01 | 61.4 | |
| SAIL-VLModel Size=2B2025.01 | 58.5 | |
| GIT_basepre-train data=10M, number of parameters=129M2023.06 | 57.5 | |
| GIT_BModel Size=0.1B2022.05 | 57.5 | |
| DocPediaResolution=2560x2560, In-house data=false2024.03 | 57.2 | |
| Qwen2-VLModel Size=8B2025.01 | 56.2 | |
| DeepSeekVL-2Model Size=8B2025.01 | 54.5 | |
| Qwen2-VLModel Size=2B2025.01 | 54.3 | |
| DeepSeekVL-2Model Size=2B2025.01 | 51.4 | |
| BLOCK+CNN+W2VOCR System=Rosetta OCR2021.12 | 48.3 | |
| Blk+CNN+W2V2023.06 | 48.3 | |
| BLOCK+CNN+W2V2022.05 | 48.3 | |
| BLOCK+CNN+W2V2022.05 | 48.3 | |
| BLOCK+CNN+W2V2022.05 | 48.3 | |
| BLOCKOCR System=Rosetta OCR2021.12 | 42 | |
| BLOCK+CNNOCR System=Rosetta OCR2021.12 | 41.5 | |
| BLIP-2Resolution=224x224, In-house data=false2024.03 | 40.6 | |
| InternVL2.5-MPOModel Size=2B2025.01 | 40 | |
| InternVL2.5-MPOModel Size=8B2025.01 | 36.7 | |
| UniDocResolution=224x224, In-house data=false2024.03 | 34.5 | |
| CNNOCR System=Rosetta OCR2021.12 | 14.3 |