Multi-modal Vision-Language Understanding on MMVet
81.3ScoreInternVL3
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| InternVL3LLM=Qwen2.5-7B, # Data=>6B / 100M / 22M, Model Scale=8B, Architecture Category=Modular, Reinforcement Learning=true2025.10 | 81.3 | — | |
| Qwen2.5-VLLLM=Qwen2.5-7B, # Data=- / - / -, Model Scale=8B, Architecture Category=Modular, Reinforcement Learning=true2025.10 | 67.1 | — | |
| InternVL2.5LLM=InternLM2.5-7B, # Data=>6B / 50M / 4M, Model Scale=8B, Architecture Category=Modular, Reinforcement Learning=false2025.10 | 62.8 | — | |
| InternVL3LLM=Qwen2.5-1.5B, # Data=>6B / 100M / 22M, Model Scale=2B, Architecture Category=Modular, Reinforcement Learning=true2025.10 | 62.2 | — | |
| Qwen2-VLLLM=Qwen2-7B, # Data=- / - / -, Model Scale=8B, Architecture Category=Modular, Reinforcement Learning=false2025.10 | 62 | — | |
| Qwen2.5-VLLLM=Qwen2.5-3B, # Data=- / - / -, Model Scale=2B, Architecture Category=Modular, Reinforcement Learning=true2025.10 | 61.8 | — | |
| InternVL2.5LLM=InternLM2.5-1.8B, # Data=>6B / 100M / 16M, Model Scale=2B, Architecture Category=Modular, Reinforcement Learning=false2025.10 | 60.8 | — | |
| Encoder-BasedLLM=Qwen3-8B, # Data=>6B / 40M / 4M, Model Scale=8B, Architecture Category=Modular, Reinforcement Learning=false2025.10 | 60 | — | |
| Mono-InternVL-1.5LLM=InternLM2-1.8B, # Data=400M / 150M / 7M, Model Scale=2B, Architecture Category=Native, Reinforcement Learning=false2025.10 | 54 | — | |
| NEOLLM=Qwen3-8B, # Data=345M / 40M / 4M, Model Scale=8B, Architecture Category=Native, Reinforcement Learning=false2025.10 | 53.6 | — | |
| NEOLLM=Qwen3-1.7B, # Data=345M / 40M / 4M, Model Scale=2B, Architecture Category=Native, Reinforcement Learning=false2025.10 | 49.6 | — | |
| Qwen2-VLLLM=Qwen2-1.5B, # Data=- / - / -, Model Scale=2B, Architecture Category=Modular, Reinforcement Learning=false2025.10 | 49.5 | — | |
| SAILLLM=Mistral-7B, # Data=512M / 86M / 6M, Model Scale=8B, Architecture Category=Native, Reinforcement Learning=false2025.10 | 46.3 | — | |
| EVEv2LLM=Qwen2.5-7B, # Data=77M / 15M / 7M, Model Scale=8B, Architecture Category=Native, Reinforcement Learning=false2025.10 | 45 | — | |
| HoVLELLM=InternLM2-1.8B, # Data=550M / 50M / 7M, Model Scale=2B, Architecture Category=Native, Reinforcement Learning=false2025.10 | 43.8 | — | |
| OneCATLLM=Qwen2.5-1.5B, # Data=436M / 70M / 13M, Model Scale=2B, Architecture Category=Native, Reinforcement Learning=false2025.10 | 42.4 | — | |
| ASPOBase Model=LLaVA-v1.5, Model Scale=13B, Optimization Protocol=ASPO2025.05 | 41.2 | 68.8 | |
| Mono-InternVLLLM=InternLM2-1.8B, # Data=1.2B / 143M / 7M, Model Scale=2B, Architecture Category=Native, Reinforcement Learning=false2025.10 | 40.1 | — | |
| BREENLLM=Qwen2.5-7B, # Data=13M / 0M / 4M, Model Scale=8B, Architecture Category=Native, Reinforcement Learning=false2025.10 | 38.9 | — | |
| DPOBase Model=LLaVA-v1.5, Model Scale=13B, Optimization Protocol=DPO2025.05 | 38.8 | 67.03 | |
| Encoder-BasedLLM=Qwen3-1.7B, # Data=>6B / 40M / 4M, Model Scale=2B, Architecture Category=Modular, Reinforcement Learning=false2025.10 | 37.4 | — | |
| Emu3LLM=from scratch, # Data=- / - / -, Model Scale=8B, Architecture Category=Native, Reinforcement Learning=false2025.10 | 37.2 | — | |
| mPLUG-Owl2-7BBase Model=mPLUG-Owl2, Model Scale=7B, Optimization Protocol=Base2025.05 | 36.2 | — | |
| LLaVA-1.5-13BBase Model=LLaVA-v1.5, Model Scale=13B, Optimization Protocol=Base2025.05 | 35.4 | 65.93 | |
| ASPOBase Model=LLaVA-v1.5, Model Scale=7B, Optimization Protocol=ASPO2025.05 | 35.3 | 65.16 | |
| VoRALLM=Qwen2.5-7B, # Data=30M / 0M / 0.6M, Model Scale=8B, Architecture Category=Native, Reinforcement Learning=false2025.10 | 33.7 | — | |
| DPOBase Model=LLaVA-v1.5, Model Scale=7B, Optimization Protocol=DPO2025.05 | 33.3 | 63.12 | |
| LLaVA-1.5-7BBase Model=LLaVA-v1.5, Model Scale=7B, Optimization Protocol=Base2025.05 | 30.5 | 62.59 | |
| SOLOLLM=Mistral-7B, # Data=44M / 0M / 2M, Model Scale=8B, Architecture Category=Native, Reinforcement Learning=false2025.10 | 30.4 | — | |
| ASPOBase Model=InstructBLIP, Model Scale=13B, Optimization Protocol=ASPO2025.05 | 27 | 45.84 | |
| LLaVA-7BBase Model=LLaVA, Model Scale=7B, Optimization Protocol=Base2025.05 | 26.7 | — | |
| DPOBase Model=InstructBLIP, Model Scale=13B, Optimization Protocol=DPO2025.05 | 26.1 | 44.5 | |
| EVELLM=Vicuna-7B, # Data=33M / 0M / 1.8M, Model Scale=8B, Architecture Category=Native, Reinforcement Learning=false2025.10 | 25.7 | — | |
| InstructBLIP-13BBase Model=InstructBLIP, Model Scale=13B, Optimization Protocol=Base2025.05 | 25.6 | 43.86 | |
| BLIP-2-7BBase Model=BLIP-2, Model Scale=7B, Optimization Protocol=Base2025.05 | 22.4 | — | |
| MiniGPT-4-7BBase Model=MiniGPT-4, Model Scale=7B, Optimization Protocol=Base2025.05 | 22.1 | — | |
| FuyuLLM=Persimmon-8B, # Data=- / - / -, Model Scale=8B, Architecture Category=Native, Reinforcement Learning=false2025.10 | 21.4 | — | |
| ChameleonLLM=from scratch, # Data=1.4B / 0M / 1.8M, Model Scale=8B, Architecture Category=Native, Reinforcement Learning=false2025.10 | 8.3 | — |