OCR and Text-based Visual Question Answering on OCRBench
84.5AccuracyQwen2 VL
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen2 VLInference mode=Zero-shot, Parameters=7B2024.10 | 84.5 | — | |
| GPT-4o-miniInference mode=Zero-shot2024.10 | 78.5 | — | |
| InternVL2 2BModel Parameters=2B, Decoding Strategy=Beam Search2024.09 | 78.1 | — | |
| GPT-4oInference mode=Zero-shot2024.10 | 73.6 | — | |
| MiniCPM-Llama3-V 2.5Inference mode=Zero-shot, Parameters=8B2024.10 | 72.5 | — | |
| Baichuan-omniInference mode=Zero-shot, Parameters=7B2024.10 | 70 | — | |
| MBQModel=InternVL2-8B, Bitwidth=W4A82024.12 | 69.6 | 73 | |
| VITAInference mode=Zero-shot, Parameters=8x7B2024.10 | 68.5 | — | |
| MBQModel=Qwen2-VL-7B, Bitwidth=W4A82024.12 | 68.4 | 72.8 | |
| MM1.5 3BModel Parameters=3B, Decoding Strategy=Greedy2024.09 | 65.7 | — | |
| Phi-3-Vision 4BModel Parameters=4B, Decoding Strategy=Greedy2024.09 | 63.7 | — | |
| MBQModel=LLaVA-onevision-7B, Bitwidth=W4A82024.12 | 63.1 | 52.3 | |
| MM1.5 1B (MoE)Model Parameters=1B, Architecture=MoE, Decoding Strategy=Greedy2024.09 | 62.6 | — | |
| MM1.5 1BModel Parameters=1B, Decoding Strategy=Greedy2024.09 | 60.5 | — | |
| MiniCPM-V2 3BModel Parameters=3B, Decoding Strategy=Beam Search2024.09 | 60.5 | — | |
| 16-bit BaselineData Format=16-bit, Backbone=LLaVA1.6-7B2024.11 | 52.4 | — | |
| AMXFP4Data Format=AMXFP4, Backbone=LLaVA1.6-7B2024.11 | 43.9 | — | |
| MXFP4Data Format=MXFP4, Backbone=LLaVA1.6-7B2024.11 | 43.4 | — | |
| MXFP4-PoTData Format=MXFP4-PoT, Backbone=LLaVA1.6-7B2024.11 | 33.7 | — | |
| AdaMMSMerging Protocol=AdaMMS2025.03 | — | 55.7 | |
| AdaMMSmerging_target=Qwen2-VL-7B, merging_source=LLaVA-OneVision-7B2025.03 | — | 85.5 | |
| CogVLM-chat-7BStatus=Base Model2025.03 | — | 56.5 | |
| DARE-LinearMerging Protocol=DARE-Linear2025.03 | — | 47.9 | |
| DARE-Linearmerging_target=Qwen2-VL-7B, merging_source=LLaVA-OneVision-7B2025.03 | — | 72.4 | |
| DARE-TiesMerging Protocol=DARE-Ties2025.03 | — | 26.5 | |
| DARE-Tiesmerging_target=Qwen2-VL-7B, merging_source=LLaVA-OneVision-7B2025.03 | — | 75.2 | |
| LLaVA-OneVisionmodel_status=Original Model2025.03 | — | 69.6 | |
| LLaVA-v1.5-7BStatus=Source Model2025.03 | — | 31.3 | |
| MetaGPTMerging Protocol=MetaGPT2025.03 | — | 56.4 | |
| MetaGPTmerging_target=Qwen2-VL-7B, merging_source=LLaVA-OneVision-7B2025.03 | — | 85.5 | |
| Qwen2-VL (base)model_status=Original Model2025.03 | — | 86 | |
| Task ArithmeticMerging Protocol=Task Arithmetic2025.03 | — | 51.2 | |
| Task Arithmeticmerging_target=Qwen2-VL-7B, merging_source=LLaVA-OneVision-7B2025.03 | — | 77.9 | |
| Ties-MergingMerging Protocol=Ties-Merging2025.03 | — | 55 | |
| Ties-Mergingmerging_target=Qwen2-VL-7B, merging_source=LLaVA-OneVision-7B2025.03 | — | 84.4 |