Scene Text Understanding on OCRBench
804ScoreInternVL2.5
Evaluation Results
| Method | Links | |
|---|---|---|
| InternVL2.5# train tokens=0.5T, Architecture Category=Token Insertion – Proprietary2025.12 | 804 | |
| Qwen2.5-VL 3B (reproduced)Res.=≤ 896²2025.12 | 804 | |
| Qwen2.5-VL 3B (reported)Res.=Native2025.12 | 797 | |
| CASA⊕ Qwen2.5-VLRes.=≤ 896²2025.12 | 790 | |
| VideoLLaMA3Architecture Category=Token Insertion – Proprietary2025.12 | 779 | |
| Qwen2-VL# train tokens=1.4T, Architecture Category=Token Insertion – Proprietary2025.12 | 767 | |
| SmolVLMArchitecture Category=Token Insertion – Public data, LLM Size=2B2025.12 | 729 | |
| InsertionHe-2B# train tokens=0.1T, Architecture Category=Token Insertion – Public data, LLM Size=2B2025.12 | 728 | |
| CASA⊕# train tokens=0.3T, Architecture=CASAHe-2B, CASA design variant=⊕, LLM Size=2B2025.12 | 728 | |
| CASA→# train tokens=0.3T, Architecture=CASAHe-2B, CASA design variant=→, LLM Size=2B2025.12 | 723 | |
| CASA∨# train tokens=0.3T, Architecture=CASAHe-2B, CASA design variant=∨, LLM Size=2B2025.12 | 694 | |
| mPLUG-Owl3 8B# train tokens=0.1T, Architecture Category=Cross-attention-based – Public data, LLM Size=8B2025.12 | 527 | |
| mPLUG-Owl3 2B# train tokens=0.1T, Architecture Category=Cross-attention-based – Public data, LLM Size=2B2025.12 | 450 | |
| InternVL3.5-38BAttack=Benign2026.07 | 0.866 | |
| Qwen2.5-VL-72B-InstructAttack=Benign2026.07 | 0.84 | |
| Qwen2.5-VL-7B-InstructAttack=Benign2026.07 | 0.837 | |
| InternVL3.5-8BAttack=Benign2026.07 | 0.827 | |
| InternVL3.5-4BAttack=Benign2026.07 | 0.813 | |
| Optimization-based Token SelectionTarget Model=InternVL3.5-38B2026.07 | 0.812 | |
| Qwen2.5-VL-3B-InstructAttack=Benign2026.07 | 0.789 | |
| Permutation AttackTarget Model=Qwen2.5-VL-3B-Instruct, Token Manipulation Ratio=0.12026.07 | 0.776 | |
| Masking AttackTarget Model=Qwen2.5-VL-3B-Instruct, Token Manipulation Ratio=0.12026.07 | 0.735 | |
| Gaussian Perturbation AttackTarget Model=Qwen2.5-VL-3B-Instruct, Token Manipulation Ratio=0.12026.07 | 0.668 | |
| Optimization-based Token SelectionTarget Model=InternVL3.5-8B2026.07 | 0.655 | |
| VTM-Attack (Naïve Average)Target Model=Qwen2.5-VL-3B-Instruct2026.07 | 0.634 | |
| Sign Flip AttackTarget Model=Qwen2.5-VL-3B-Instruct, Token Manipulation Ratio=0.12026.07 | 0.357 | |
| Optimization-based Token SelectionTarget Model=Qwen2.5-VL-3B-Instruct, Base Attack=Sign Flip, Iterations=30, Learning Rate=0.052026.07 | 0.305 | |
| Optimization-based Token SelectionTarget Model=InternVL3.5-4B2026.07 | 0.157 | |
| Optimization-based Token SelectionTarget Model=Qwen2.5-VL-7B-Instruct2026.07 | 0.15 | |
| Optimization-based Token SelectionTarget Model=Qwen2.5-VL-72B-Instruct2026.07 | 0.02 |