Visual Question Answering on OCRVQA
87.5AccuracyOracle
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| OracleData Source Type=Real Data Reweighting, Sampling Strategy Category=Gold, Compute Budget=280K datapoints2026.03 | 87.5 | — | |
| D.P.Data Source Type=Real Data Reweighting, Sampling Strategy Category=Ours, Compute Budget=280K datapoints2026.03 | 87.4 | — | |
| ICONSData Source Type=Real Data Reweighting, Sampling Strategy Category=Baseline, Compute Budget=280K datapoints2026.03 | 86.5 | — | |
| UniformData Source Type=Real Data Reweighting, Sampling Strategy Category=Baseline, Compute Budget=280K datapoints2026.03 | 85.2 | — | |
| Cont-Squeeze (128 -> 1)Optimization steps=+700 steps2026.02 | 84.86 | — | |
| In-Squeeze (128 ... -> 1)Optimization mode=Min steps2026.02 | 83.93 | — | |
| Cont-Squeeze (128 -> 1)Optimization steps=+200 steps2026.02 | 83.82 | — | |
| In-Squeeze (128 ... -> 1)Optimization mode=Standard2026.02 | 83.79 | — | |
| Direct Fine-tuning (r=1)Optimization steps=+700 steps2026.02 | 83.66 | — | |
| Direct Fine-tuning (r=1)Optimization steps=+0 steps2026.02 | 83.14 | — | |
| Direct Fine-tuning (r=1)Optimization steps=+200 steps2026.02 | 83.06 | — | |
| Cont-Squeeze (128 -> 1)Optimization steps=+0 steps2026.02 | 82.73 | — | |
| Qwen-VLModel Category=generalist models2023.12 | 75.7 | — | |
| PALI-X-55BModel Category=task-specific fine-tuning models2023.12 | 75 | — | |
| CogAgentModel Category=generalist models2023.12 | 75 | — | |
| D.P.Data Source Type=Synthetic Data Selection, Sampling Strategy Category=Ours, Compute Budget=280K datapoints2026.03 | 74.6 | — | |
| Qwen2.5-VL-3BNumber of parameters=3B2026.01 | 74.53 | — | |
| CogVLMModel Category=task-specific fine-tuning models2023.12 | 74.5 | — | |
| Qwen2-VL-2BNumber of parameters=2B2026.01 | 74.17 | — | |
| CogVLMModel Category=generalist models2023.12 | 74.1 | — | |
| ICONSData Source Type=Synthetic Data Selection, Sampling Strategy Category=Baseline, Compute Budget=280K datapoints2026.03 | 73.5 | — | |
| BLIP-2Model Category=task-specific fine-tuning models2023.12 | 72.7 | — | |
| LoRAModel=Qwen2.5-VL-3B, Optimization=FlatPO2026.06 | 71.4 | — | |
| LoRAModel=Qwen2.5-VL-3B, Optimization=SAM2026.06 | 70.8 | — | |
| Qwen-VL-chatModel Category=generalist models2023.12 | 70.5 | — | |
| LoRAModel=Qwen2.5-VL-3B, Optimization=GAM2026.06 | 70.5 | — | |
| Prefix TuningModel=Qwen2.5-VL-3B, Optimization=FlatPO2026.06 | 69.3 | — | |
| LoRAModel=Qwen2.5-VL-3B, Optimization=Base2026.06 | 68.6 | — | |
| Prefix TuningModel=Qwen2.5-VL-3B, Optimization=GAM2026.06 | 68.3 | — | |
| Prefix TuningModel=Qwen2.5-VL-3B, Optimization=SAM2026.06 | 68.2 | — | |
| AuroraEdge-V-2BNumber of parameters=2B2026.01 | 68.15 | — | |
| Prefix TuningModel=Qwen2.5-VL-3B, Optimization=Base2026.06 | 67.2 | — | |
| UniformData Source Type=Synthetic Data Selection, Sampling Strategy Category=Baseline, Compute Budget=280K datapoints2026.03 | 65.4 | — | |
| Gemma 3 4B ITZero-shot=true2026.02 | 65.3 | — | |
| LLaVA1.5Evaluation Protocol=Zero-shot2024.06 | 58.1 | — | |
| VisionTrimBackbone=LLaVA-1.5-13B, Tokens=322026.01 | 55.2 | 704 | |
| VanillaBackbone=LLaVA-1.5-13B, Tokens=322026.01 | 54.9 | 2,394 | |
| LLaVAROCR Usage=w/ OCR2024.06 | 42.7 | — | |
| VisionZipBackbone=LLaVA-1.5-13B, Tokens=322026.01 | 42.3 | 1,096 | |
| LaRAEvaluation Protocol=Zero-shot2024.06 | 41.2 | — | |
| InternVL-2.5-2BNumber of parameters=2B2026.01 | 32.02 | — | |
| BLIP-2Evaluation Protocol=Zero-shot2024.06 | 30.7 | — | |
| LLaVAREvaluation Protocol=Finetuned2024.06 | 28.8 | — | |
| mPLUG-OwlEvaluation Protocol=Zero-shot2024.06 | 28.6 | — | |
| mPLUG-Owl2Evaluation Protocol=Zero-shot2024.06 | 28.6 | — | |
| OpenFlamingoEvaluation Protocol=Zero-shot2024.06 | 27.8 | — | |
| LLaVAREvaluation Protocol=Zero-shot2024.06 | 23.8 | — | |
| Original ModelBase Model=InternVL-Chat-ViT-6B-Vicuna-13B2025.12 | 16.5 | — | |
| TTAdaptstrategy=Model parameter adaptation (2)2025.10 | 13.8 | — | |
| TrimTokenator-LCBase Model=InternVL-Chat-ViT-6B-Vicuna-13B2025.12 | 13 | — | |
| TTAugstrategy=Test-Time Augmentation2025.10 | 12.6 | — | |
| DivPruneBase Model=InternVL-Chat-ViT-6B-Vicuna-13B2025.12 | 12.5 | — | |
| CATPBase Model=InternVL-Chat-ViT-6B-Vicuna-13B2025.12 | 12 | — | |
| TTAdaptstrategy=Model parameter adaptation (1)2025.10 | 11.9 | — | |
| VTWBase Model=InternVL-Chat-ViT-6B-Vicuna-13B2025.12 | 11.5 | — | |
| VisionZipBase Model=InternVL-Chat-ViT-6B-Vicuna-13B2025.12 | 11.5 | — | |
| TrimTokenatorBase Model=InternVL-Chat-ViT-6B-Vicuna-13B2025.12 | 11.5 | — | |
| MiniGPT4Evaluation Protocol=Zero-shot2024.06 | 11.5 | — | |
| SparseVLMBase Model=InternVL-Chat-ViT-6B-Vicuna-13B2025.12 | 11 | — | |
| LLaVAEvaluation Protocol=Zero-shot2024.06 | 11 | — | |
| FastVBase Model=InternVL-Chat-ViT-6B-Vicuna-13B2025.12 | 10.5 | — | |
| Baselinestrategy=Baseline2025.10 | 0 | — |