Text-Centric Vision-Language Understanding on OCR Bench
517AccuracyInternVL
Evaluation Results
| Method | Links | |
|---|---|---|
| InternVLModel Generation Capability=Text only2024.07 | 517 | |
| MonkeyModel Generation Capability=Text only2024.07 | 514 | |
| InternLM-XComposer2Model Generation Capability=Text only2024.07 | 511 | |
| TextHarmony-ChatModel Generation Capability=Text and image, Slide-LoRA=true2024.07 | 448 | |
| TextHarmonyModel Generation Capability=Text and image, Slide-LoRA=true2024.07 | 440 | |
| TextHarmony*Model Generation Capability=Text and image, Slide-LoRA=false2024.07 | 397 | |
| mPLUG-Owl2Model Generation Capability=Text only2024.07 | 366 | |
| SEED-LLaMA-14BModel Generation Capability=Text and image2024.07 | 357 | |
| LLaVARModel Generation Capability=Text only2024.07 | 346 | |
| LLaVA1.5-7BModel Generation Capability=Text only2024.07 | 297 | |
| MM-InterleavedModel Generation Capability=Text and image2024.07 | 197 | |
| Qwen2.5-VL-3B-InstructParams=3.75B2025.12 | 82.7 | |
| Qwen2-VL-2BParams=2.2B2025.12 | 80.9 | |
| InternVL2-2BParams=2.2B2025.12 | 78.4 | |
| InternVL2-1BParams=0.94B2025.12 | 75.7 | |
| MobileNet-QwenParams=1.84B2025.12 | 73.5 | |
| InternViT-QwenParams=1.86B2025.12 | 72.5 | |
| AutoNeuralParams=1.47B2025.12 | 71.4 | |
| MiniGPT5Model Generation Capability=Text and image2024.07 | 68 | |
| InternViT-LiquidParams=1.49B2025.12 | 66 |