Information Visual Question Answering on InfoVQA (test)
89.3ANLSSeed-1.5-VL
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Seed-1.5-VLSize=20B2026.03 | 89.3 | — | |
| Pixelis (Qwen3-VL-8B-Instruct)Size=8B2026.03 | 87.9 | — | |
| Pixelis: SFT + RFTSize=8B2026.03 | 86.7 | — | |
| PRM (process reward; tools; 8B)Size=8B2026.03 | 86.2 | — | |
| Pixelis: SFT + TTRLSize=8B2026.03 | 86.1 | — | |
| Pixel Reasoner (Qwen3-VL-8B-Instruct)Size=8B2026.03 | 85.8 | — | |
| Pixelis: RFT + TTRLSize=8B2026.03 | 85.8 | — | |
| Late-fusionSize=8B2026.03 | 85.6 | — | |
| Step Self-Consistency (step-level)Size=8B2026.03 | 85.4 | — | |
| Pixelis: SFT onlySize=8B2026.03 | 84.9 | — | |
| Qwen-VL-Max-0809Params (B)=722024.10 | 84.5 | — | |
| Qwen2-VL-72B2024.12 | 84.5 | — | |
| Pixelis: RFT onlySize=8B2026.03 | 84.4 | — | |
| InternVL2.5-78B2024.12 | 84.1 | — | |
| RA-TTA (retrieval-augmented)Size=8B2026.03 | 84 | — | |
| InternVL2.5-38B2024.12 | 83.6 | — | |
| RV Self-Consistency (answer-only)Size=8B2026.03 | 83.6 | — | |
| Pixelis: TTRL onlySize=8B2026.03 | 83.5 | — | |
| Qwen3-VL-8B-InstructSize=8B2026.03 | 83.1 | — | |
| Realistic TTA of VLMs (StatA)Size=8B2026.03 | 82.7 | — | |
| InternVL2-Llama3-76B2024.12 | 82 | — | |
| Qwen3-VL-30B-A3B-InstructSize=30B2026.03 | 82 | — | |
| Molmo-72B2024.12 | 81.9 | — | |
| Gemini-1.5-ProModel Scale=High-End2024.09 | 81 | — | |
| Gemini-1.5-Pro2024.12 | 81 | — | |
| SmoLA PaLI-XOCR Usage=SOTA Overall2024.02 | 80.3 | — | |
| InternVL2.5-26B2024.12 | 79.8 | — | |
| GPT-4o-202405132024.12 | 79.2 | — | |
| InternVL2-40B2024.12 | 78.7 | — | |
| DocCogitoSize(B)=82026.03 | 78.6 | — | |
| InternVL2.5-8B2024.12 | 77.6 | — | |
| Qwen2-VL-7B2024.12 | 76.5 | — | |
| DocCogitoSize(B)=42026.03 | 76 | — | |
| InternVL2-26B2024.12 | 75.9 | — | |
| InternVL-1.5-Plusparameters=26B2024.11 | 75.7 | — | |
| Qwen3-VL-InstructSize(B)=82026.03 | 75.7 | — | |
| MartenSize(B)=8.12026.03 | 75.2 | — | |
| GPT-4V2024.12 | 75.1 | — | |
| LLaVA-OneVision-72B2024.12 | 74.9 | — | |
| InternVL2-8B2024.12 | 74.8 | — | |
| InternVL2Size(B)=8.12026.03 | 74.8 | — | |
| Qwen3-VL-InstructSize(B)=42026.03 | 74.6 | — | |
| Claude-3.5-Sonnet2024.12 | 74.3 | — | |
| Molmo-7B-D2024.12 | 72.6 | — | |
| InternVL-Chat-V1.52024.12 | 72.5 | — | |
| InternVL2.5-4B2024.12 | 72.1 | — | |
| TextHawk2Size(B)=7.42026.03 | 67.8 | — | |
| MM1.5-30BModel Scale=30B2024.09 | 67.3 | — | |
| InternVL2-4B2024.12 | 67 | — | |
| SMoLA-PaLI-X_FT (Specialist)Base model=PaLI-X, Per-task LoRA tuning=true, LoRA rank=42023.12 | 66.2 | — | |
| ScreenAIOCR Usage=With OCR2024.02 | 65.9 | — | |
| SMoLA-PaLI-XFTOCR pipeline input=true, Model Category=Generalist, Backbone=PaLI-X2023.12 | 65.6 | — | |
| Qwen2-VL-2B2024.12 | 65.5 | — | |
| PaLI-3 (Specialist)OCR pipeline input=true, Model Category=Specialist, Reference Citation=[8]2023.12 | 62.4 | — | |
| SOTA [8]2023.12 | 62.4 | — | |
| PaLI-3OCR Usage=With OCR2024.02 | 62.4 | — | |
| ScreenAIOCR Usage=Without OCR2024.02 | 61.4 | — | |
| InternVL2.5-2B2024.12 | 60.9 | — | |
| Mini-MonkeySize(B)=22026.03 | 60.1 | — | |
| MM1.5-7BModel Scale=7B2024.09 | 59.5 | — | |
| InternVL2-2BModel Scale=1B2024.09 | 58.9 | — | |
| InternVL2-2B2024.12 | 58.9 | — | |
| MM1.5-3BModel Scale=3B2024.09 | 58.5 | — | |
| DocLayLLMSize(B)=82026.03 | 58.4 | — | |
| Aquila-VL-2B2024.12 | 58.3 | — | |
| PaLI-3 (Specialist)OCR pipeline input=false, Model Category=Specialist, Reference Citation=[8]2023.12 | 57.8 | — | |
| PaLI-3OCR Usage=Without OCR2024.02 | 57.8 | — | |
| SMoLA-PaLI-3FTOCR pipeline input=true, Model Category=Generalist, Backbone=PaLI-32023.12 | 57.3 | — | |
| InternVL2.5-1B2024.12 | 56 | — | |
| MM1.5-1B-MOEModel Scale=1B, Architecture=MoE2024.09 | 55.9 | — | |
| Claude-3-Opus2024.12 | 55.6 | — | |
| PaLI-XBase model=PaLI-X2023.12 | 54.8 | — | |
| Gemini Nano-2Model Scale=3B2024.09 | 54.5 | — | |
| MM1.5-3B-MOEModel Scale=3B, Architecture=MoE2024.09 | 53.6 | — | |
| SMoLA-PaLI-3FTOCR pipeline input=false, Model Category=Generalist, Backbone=PaLI-32023.12 | 52.4 | — | |
| Gemini Nano-1Model Scale=1B2024.09 | 51.1 | — | |
| InternVL2-1B2024.12 | 50.9 | — | |
| DocOwl-1.5-ChatModel Scale=7B2024.09 | 50.7 | — | |
| DocOwl-1.5-ChatSize(B)=8.12026.03 | 50.7 | — | |
| MM1.5-1BModel Scale=1B2024.09 | 50.5 | — | |
| DocOwl-1.5Size(B)=8.12026.03 | 50.4 | — | |
| Zhang et al.Size(B)=8.12026.03 | 50.2 | — | |
| SMoLA-PaLI-XFTOCR pipeline input=false, Model Category=Generalist, Backbone=PaLI-X2023.12 | 49.2 | — | |
| Phi-3-Vision-4BModel Scale=4B2024.09 | 49 | — | |
| DocFormerv2largeModality=image + text + spatial features, Pre-train data=64M, #param=750M, Extra document VQA data=true2023.06 | 48.8 | — | |
| DocFormerv2Size(B)=0.752026.03 | 48.8 | — | |
| UDOPModality=image + text + spatial features, Pre-train data=11M, #param=794M2023.06 | 47.4 | — | |
| UDOPtype=specialized pipeline2023.05 | 47.4 | — | |
| MM1-30BModel Scale=30B2024.09 | 47.3 | — | |
| DocKylinSize(B)=7.12026.03 | 46.6 | — | |
| DocOwl2Size(B)=82026.03 | 46.4 | — | |
| T5large+2D+UModality=text / (text + spatial) features, #param=750M2023.06 | 46.1 | — | |
| Cambrian-34B2024.12 | 46 | — | |
| HVFAbackbone=mPLUG-Owl, fine-tuning dataset=650K2024.11 | 45.9 | — | |
| mPLUG-Owl-7B + OursModel Params=7.2B, Trainable Params=96M, Pre-training Data=1.1B, Fine-tuning Data=650K2024.11 | 45.9 | — | |
| Park et al.Size(B)=7.22026.03 | 45.9 | — | |
| MM1-7BModel Scale=7B2024.09 | 45.5 | — | |
| LayoutLMv3largeModality=image + text + spatial features, Pre-train data=11M, #param=368M2023.06 | 45.1 | — | |
| MM1-3BModel Scale=3B2024.09 | 44.7 | — | |
| CogAgentSize(B)=17.32026.03 | 44.5 | — |