Multi-modal Evaluation on MME (total)
2,433.61MME Total ScoreQwen-VL-Max
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen-VL-Max2024.05 | 2,433.61 | |
| GPT-4oindependent reproduction=true, IoT prompting=true2024.05 | 2,396.1 | |
| GPT-4oindependent reproduction=true, IoT prompting=false2024.05 | 2,254.1 | |
| InternLM-XComposer2-VLParams=7B2024.05 | 2,242.71 | |
| Qwen-VL-Plus2024.05 | 2,183.39 | |
| NextFlow#Training Data=40M, Model Size=7B2026.01 | 1,897.7 | |
| LLaVA 1.5#Training Data=0.7M, Model Size=13B2026.01 | 1,826.7 | |
| NextFlow#Training Data=0.7M, Model Size=7B2026.01 | 1,752.1 | |
| NextFlow#Training Data=21M, Model Size=7B2026.01 | 1,744.9 | |
| MAPBackbone=InternVL32025.08 | 1,737 | |
| VanillaBackbone=InternVL32025.08 | 1,729.6 | |
| DCLABackbone=InternVL32025.08 | 1,726.4 | |
| DAMOBackbone=InternVL32025.08 | 1,725.8 | |
| SPINBackbone=InternVL32025.08 | 1,722.8 | |
| MAPBackbone=QwenVL2.52025.08 | 1,712 | |
| DAMOBackbone=InternVL2.52025.08 | 1,709.9 | |
| MAPBackbone=InternVL2.52025.08 | 1,704.3 | |
| VanillaBackbone=QwenVL2.52025.08 | 1,704.2 | |
| DCLABackbone=QwenVL2.52025.08 | 1,704 | |
| DAMOBackbone=QwenVL2.52025.08 | 1,699.9 | |
| VanillaBackbone=InternVL2.52025.08 | 1,697.2 | |
| DCLABackbone=InternVL2.52025.08 | 1,694.6 | |
| SPINBackbone=QwenVL2.52025.08 | 1,686.3 | |
| Gemini-Pro-1.5independent reproduction=true, IoT prompting=true2024.05 | 1,686.1 | |
| SPINBackbone=InternVL2.52025.08 | 1,685.2 | |
| LLaVA-v1.5-13BParams=13.4B2024.05 | 1,615.6 | |
| Gemini-Pro-1.5independent reproduction=true, IoT prompting=false2024.05 | 1,571.4 | |
| LLaVA-1.5-7BReduction Ratio=0%, Vision Tokens=5762026.02 | 1,510.7 | |
| MQTReduction Ratio=88%, Vision Tokens=642026.02 | 1,464.3 | |
| M3Reduction Ratio=75%, Vision Tokens=362026.02 | 1,417.2 | |
| MQTReduction Ratio=88%, Vision Tokens=362026.02 | 1,416.3 | |
| Mask-LLaVAReduction Ratio=90%, Vision Tokens=572026.02 | 1,415 | |
| MQTReduction Ratio=88%, Vision Tokens=162026.02 | 1,408.5 | |
| Mask-LLaVAReduction Ratio=97%, Vision Tokens=422026.02 | 1,402.7 | |
| Mask-LLaVAReduction Ratio=97%, Vision Tokens=152026.02 | 1,395.8 | |
| M3Reduction Ratio=75%, Vision Tokens=92026.02 | 1,374.8 | |
| FasterVLMReduction Ratio=90%, Vision Tokens=582026.02 | 1,348.6 | |
| FasterVLMReduction Ratio=95%, Vision Tokens=292026.02 | 1,254.8 | |
| LLaVA-1.5-7B (Random Patch Dropping)Reduction Ratio=90%, Vision Tokens=582026.02 | 1,246.8 | |
| FitPruneReduction Ratio=90%, Vision Tokens=582026.02 | 1,147.4 | |
| LLaVA-1.5-7B (Random Patch Dropping)Reduction Ratio=95%, Vision Tokens=292026.02 | 1,141.1 | |
| SparseVLMReduction Ratio=90%, Vision Tokens=582026.02 | 1,030.6 | |
| FitPruneReduction Ratio=95%, Vision Tokens=292026.02 | 855.2 | |
| MiniGPT-4Params=8B2024.05 | 725.95 | |
| AdaMMSmerging_target=Qwen2-VL-7B, merging_source=LLaVA-OneVision-7B2025.03 | 83.36 | |
| Ties-Mergingmerging_target=Qwen2-VL-7B, merging_source=LLaVA-OneVision-7B2025.03 | 82.65 | |
| Task Arithmeticmerging_target=Qwen2-VL-7B, merging_source=LLaVA-OneVision-7B2025.03 | 82.33 | |
| Qwen2-VL (base)model_status=Original Model2025.03 | 81.44 | |
| MetaGPTmerging_target=Qwen2-VL-7B, merging_source=LLaVA-OneVision-7B2025.03 | 81.21 | |
| LLaVA-OneVisionmodel_status=Original Model2025.03 | 77.04 | |
| AdaMMSMerging Protocol=AdaMMS2025.03 | 69.09 | |
| LLaVA-v1.5-7BStatus=Source Model2025.03 | 66.68 | |
| LLaVAMethod Category=Original Model2025.03 | 66.68 | |
| DARE-Linearmerging_target=Qwen2-VL-7B, merging_source=LLaVA-OneVision-7B2025.03 | 66.06 | |
| Task ArithmeticMerging Protocol=Task Arithmetic2025.03 | 65.99 | |
| Task Arithmeticunsupervised=false2025.03 | 64.65 | |
| AdaMMSunsupervised=true2025.03 | 64.65 | |
| AdaMMSUnsupervised=true2025.03 | 64.61 | |
| MetaGPTUnsupervised=true2025.03 | 64.24 | |
| Ties-MergingUnsupervised=false2025.03 | 64.2 | |
| DARE-LinearMerging Protocol=DARE-Linear2025.03 | 64.08 | |
| Task ArithmeticUnsupervised=false2025.03 | 63.17 | |
| DARE-LinearUnsupervised=false2025.03 | 62.99 | |
| mPLUG-Owl2(base)2025.03 | 62.8 | |
| mPLUG-Owl2(base)Method Category=Original Model2025.03 | 62.8 | |
| DARE-Linearunsupervised=false2025.03 | 62.44 | |
| DARE-TiesUnsupervised=false2025.03 | 60.37 | |
| MetaGPTMerging Protocol=MetaGPT2025.03 | 59.37 | |
| CogVLM-chat-7BStatus=Base Model2025.03 | 59.23 | |
| CogVLM2025.03 | 59.23 | |
| DARE-Tiesunsupervised=false2025.03 | 57.9 | |
| Ties-MergingMerging Protocol=Ties-Merging2025.03 | 57.29 | |
| MetaGPTunsupervised=true2025.03 | 56.81 | |
| DARE-Tiesmerging_target=Qwen2-VL-7B, merging_source=LLaVA-OneVision-7B2025.03 | 54.43 | |
| Ties-Mergingunsupervised=false2025.03 | 48.96 | |
| DARE-TiesMerging Protocol=DARE-Ties2025.03 | 46.75 |