Multimodal Reasoning on MM-Vet
86.2MM-Vet ScoreGenRecal (InternVL3.5-8B)
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| GenRecal (InternVL3.5-8B)Teacher VLM=InternVL3.5-38B2025.06 | 86.2 | — | — | — | — | — | — | — | 83.4 | |
| GenRecal (InternVL3.5-8B)Teacher VLM=Qwen3-VL-32B2025.06 | 86 | — | — | — | — | — | — | — | 83.4 | |
| InternVL3.5-8B2025.06 | 83.1 | — | — | — | — | — | — | — | 80.4 | |
| InternVL3.5Size=8B, mode=instruct2026.04 | 83.1 | — | — | — | — | — | — | — | — | |
| InternVL3.5-38B2025.06 | 82.2 | — | — | — | — | — | — | — | 82.8 | |
| Gemini 2.5 FlashSize=-, mode=instruct2026.04 | 81.4 | — | — | — | — | — | — | — | — | |
| GPT-4o-0806Model Source=Close-Source2025.01 | 80.8 | — | — | — | — | — | — | — | — | |
| Qwen3-VL-32B2025.06 | 79.4 | — | — | — | — | — | — | — | 82.5 | |
| GenRecal (Qwen3-VL-8B)Teacher VLM=Qwen3-VL-32B2025.06 | 77.6 | — | — | — | — | — | — | — | 80.1 | |
| GenRecal (Qwen3-VL-8B)Teacher VLM=InternVL3.5-38B2025.06 | 77.4 | — | — | — | — | — | — | — | 80 | |
| Qwen3-OmniSize=30B-A3B, mode=instruct2026.04 | 74.8 | — | — | — | — | — | — | — | — | |
| GPT-4o-mini-0718Model Source=Close-Source2025.01 | 74.6 | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8B2025.06 | 74.5 | — | — | — | — | — | — | — | 77 | |
| MiniCPM-o 4.5Size=9B, mode=instruct2026.04 | 74.4 | — | — | — | — | — | — | — | — | |
| Llama-3.2-90B-Vision-InstModel Source=Open-Source2025.01 | 74.1 | — | — | — | — | — | — | — | — | |
| Qwen2-VL-72B2025.06 | 74 | — | — | — | — | — | — | — | — | |
| Qwen3-VLSize=8B, mode=instruct2026.04 | 73.7 | — | — | — | — | — | — | — | — | |
| PG-CoTParameters=7B2025.12 | 73.5 | — | — | — | — | — | — | — | — | |
| InternVL2.5-8B-GenRecalTeacher VLM=InternVL2.5-78B2025.06 | 73.2 | — | — | — | — | — | — | — | — | |
| InternVL2.5-8B-GenRecalTeacher VLM=InternVL2-76B2025.06 | 72.4 | — | — | — | — | — | — | — | — | |
| InternVL2.5-78B2025.06 | 72.3 | — | — | — | — | — | — | — | — | |
| InternVL2.5-8B-GenRecalTeacher VLM=Qwen2-VL-72B2025.06 | 71.4 | — | — | — | — | — | — | — | — | |
| Gemini-1.5-ProModel Source=Close-Source2025.01 | 71.3 | — | — | — | — | — | — | — | — | |
| PhantomParameters=7B2024.09 | 70.8 | — | — | — | — | — | — | — | — | |
| VanillaBackbone=InternVL3-8B, Pruning Stage=N/A, Token Retention Rate=100%2026.01 | 70.41 | — | — | — | — | — | — | — | — | |
| Claude-3.5-Sonnet2025.06 | 70.1 | — | — | — | — | — | — | — | — | |
| CAPABackbone=InternVL3-8B, Pruning Stage=Transition, Token Retention Rate=25%2026.01 | 69.85 | — | — | — | — | — | — | — | — | |
| AD-Loop#Params=7B, Capabilities=Und. and Gen.2026.02 | 69.7 | — | — | — | — | — | — | — | — | |
| EVE-iter3Backbone=MiMo-VL-7B-SFT-2508, Training Iteration=32026.04 | 69.63 | — | — | — | — | — | — | — | — | |
| EVE (ours-iter3)Backbone=Qwen3-VL-4B, Iteration=32026.04 | 69.36 | — | — | — | — | — | — | — | — | |
| GPT-4o (0513)2025.06 | 69.1 | — | — | — | — | — | — | — | — | |
| EVE (ours-iter1)Backbone=Qwen3-VL-4B, Iteration=12026.04 | 68.99 | — | — | — | — | — | — | — | — | |
| Claude3.5-Sonnet-0620Model Source=Close-Source2025.01 | 68.7 | — | — | — | — | — | — | — | — | |
| EVE (ours-iter2)Backbone=Qwen3-VL-4B, Iteration=22026.04 | 68.58 | — | — | — | — | — | — | — | — | |
| EVE-iter1Backbone=MiMo-VL-7B-SFT-2508, Training Iteration=12026.04 | 68.53 | — | — | — | — | — | — | — | — | |
| MMRPTModel Scale=Qwen2.5-VL-7B-Instruct, Evaluation Protocol=Zero-shot2025.12 | 68.12 | — | — | — | — | — | — | — | — | |
| Qwen3-VL-4B-InstructBackbone=Qwen3-VL-4B, Iteration=Base2026.04 | 67.89 | — | — | — | — | — | — | — | — | |
| ThinkLite-VLParameters=7B2025.12 | 67.8 | — | — | — | — | — | — | — | — | |
| EVE-iter2Backbone=MiMo-VL-7B-SFT-2508, Training Iteration=22026.04 | 67.52 | — | — | — | — | — | — | — | — | |
| BaseModel Scale=Qwen2.5-VL-7B-Instruct, Evaluation Protocol=Zero-shot2025.12 | 67.4 | — | — | — | — | — | — | — | — | |
| BAGEL2025.06 | 67.2 | — | — | — | — | — | — | — | — | |
| BAGEL#Params=7B, Capabilities=Und. and Gen.2026.02 | 67.2 | — | — | — | — | — | — | — | — | |
| BAGEL#Params=7B MoT, Model Category=Unified Multimodal Large Language Models, Co-training=false2026.05 | 67.2 | — | — | — | — | — | — | — | — | |
| UniWorld-V12025.06 | 67.1 | — | — | — | — | — | — | — | — | |
| Qwen2.5-VLParameters=7B2025.12 | 67.1 | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL#Params=7B, Model Category=Vision Language Models, Co-training=false2026.05 | 67.1 | — | — | — | — | — | — | — | — | |
| MIRRORReasoning Paradigm=Thinking with Images, Base Model=Qwen2.5-VL-7B2026.02 | 66.7 | — | — | — | — | — | — | — | — | |
| Kimi-VL#Params=3B/16B, Model Category=Vision Language Models, Co-training=false2026.05 | 66.7 | — | — | — | — | — | — | — | — | |
| MiMo-VL-7B-SFT-2508Backbone=MiMo-VL-7B-SFT-2508, Training Iteration=02026.04 | 66.67 | — | — | — | — | — | — | — | — | |
| MetaQuery-XL2025.06 | 66.6 | — | — | — | — | — | — | — | — | |
| Adaptive-CoFReasoning Paradigm=Thinking with Images, Base Model=Qwen2.5-VL-7B2026.02 | 66.21 | — | — | — | — | — | — | — | — | |
| InternVL2-76B2025.06 | 65.7 | — | — | — | — | — | — | — | — | |
| LlamaV-o1Model Source=Our Models2025.01 | 65.4 | — | — | — | — | — | — | — | — | |
| Intern3.5-VLLanguage Backbone=Qwen3 1.7B LLM, Evaluation Framework=VLMEvalKit [69]2026.03 | 65.2 | — | — | — | — | — | — | — | — | |
| CAPABackbone=Qwen2.5-VL-7B, Pruning Stage=Transition, Token Retention Rate=25%2026.01 | 64.25 | — | — | — | — | — | — | — | — | |
| Gemini-1.5-Pro2025.06 | 64 | — | — | — | — | — | — | — | — | |
| InternVL2.5-8B-GenRecalTeacher VLM=NVLM-72B2025.06 | 63.9 | — | — | — | — | — | — | — | — | |
| GPT-4V2024.09 | 63.6 | — | — | — | — | — | — | — | — | |
| UAM#Params=7B MoT, Model Category=Vision-Language-Action Models, Co-training=false2026.05 | 63.4 | — | — | — | — | — | — | — | — | |
| Qwen2-VLParameters=7B2024.09 | 62 | — | — | — | — | — | — | — | — | |
| Qwen2-VL#Params=7B, Model Category=Vision Language Models, Co-training=false2026.05 | 62 | — | — | — | — | — | — | — | — | |
| VanillaBackbone=Qwen2.5-VL-7B, Pruning Stage=N/A, Token Retention Rate=100%2026.01 | 61.83 | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL#Params=3B, Model Category=Vision Language Models, Co-training=false2026.05 | 61.8 | — | — | — | — | — | — | — | — | |
| Molmo-72B2025.06 | 61.1 | — | — | — | — | — | — | — | — | |
| MMRPTModel Scale=Qwen2.5-VL-3B-Instruct, Evaluation Protocol=Zero-shot2025.12 | 60.62 | — | — | — | — | — | — | — | — | |
| LLaVA-OneVision-72B2025.06 | 60.6 | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7B + SRPOBackbone=Qwen2.5-VL-7B2026.05 | 60.57 | — | — | — | — | — | — | — | — | |
| Llava-CoTModel Source=Open-Source2025.01 | 60.3 | — | — | — | — | — | — | — | — | |
| DeepEyesReasoning Paradigm=Thinking with Images, Base Model=Qwen2.5-VL-7B2026.02 | 60.28 | — | — | — | — | — | — | — | — | |
| Qwen3-VLLanguage Backbone=Qwen3 1.7B LLM, Evaluation Framework=VLMEvalKit [69]2026.03 | 60.18 | — | — | — | — | — | — | — | — | |
| MiniCPM-V-2.6Parameters=7B2024.09 | 60 | — | — | — | — | — | — | — | — | |
| InternVL2Parameters=8B2024.09 | 60 | — | — | — | — | — | — | — | — | |
| DeepSeek-VL2#Params=4B/27B, Model Category=Vision Language Models, Co-training=false2026.05 | 60 | — | — | — | — | — | — | — | — | |
| BaseModel Scale=Qwen2.5-VL-3B-Instruct, Evaluation Protocol=Zero-shot2025.12 | 59.4 | — | — | — | — | — | — | — | — | |
| NVLM-72B2025.06 | 58.9 | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7B + DAPOBackbone=Qwen2.5-VL-7B2026.05 | 58.02 | — | — | — | — | — | — | — | — | |
| Vision-SR1-7BBackbone=Qwen2.5-VL-7B2026.05 | 57.98 | — | — | — | — | — | — | — | — | |
| PAPO-G-7BBackbone=Qwen2.5-VL-7B2026.05 | 57.93 | — | — | — | — | — | — | — | — | |
| PAPO-D-7BBackbone=Qwen2.5-VL-7B2026.05 | 57.66 | — | — | — | — | — | — | — | — | |
| Llama-3.2-11B-Vision-InstModel Source=Our Models, Note=baseline2025.01 | 57.6 | — | — | — | — | — | — | — | — | |
| LLaVA-OV2025.06 | 57.5 | — | — | — | — | — | — | — | — | |
| LLaVA-OneVisionParameters=8B2024.09 | 57.5 | — | — | — | — | — | — | — | — | |
| LLava-OV#Params=7B, Model Category=Vision Language Models, Co-training=false2026.05 | 57.5 | — | — | — | — | — | — | — | — | |
| Vision-Matters-7BBackbone=Qwen2.5-VL-7B2026.05 | 57.43 | — | — | — | — | — | — | — | — | |
| LLaVA-NeXT2025.06 | 57.4 | — | — | — | — | — | — | — | — | |
| Meteor2024.05 | 57.3 | — | — | — | — | — | — | — | — | |
| Meteor2024.05 | 57.3 | — | — | — | — | — | — | — | — | |
| Perception-R1-7BBackbone=Qwen2.5-VL-7B2026.05 | 57.24 | — | — | — | — | — | — | — | — | |
| InternVL2-8BModel Source=Open-Source2025.01 | 56.9 | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7B + GRPOBackbone=Qwen2.5-VL-7B2026.05 | 56.81 | — | — | — | — | — | — | — | — | |
| MiniCPM-V2.6-8BModel Source=Open-Source2025.01 | 56.3 | — | — | — | — | — | — | — | — | |
| VL-Rethinker-7BBackbone=Qwen2.5-VL-7B2026.05 | 56.23 | — | — | — | — | — | — | — | — | |
| VL-RethinkerReasoning Paradigm=Text Reflection, Base Model=Qwen2.5-VL-7B2026.02 | 56.19 | — | — | — | — | — | — | — | — | |
| VPPO-7BBackbone=Qwen2.5-VL-7B2026.05 | 56.01 | — | — | — | — | — | — | — | — | |
| PRCO-7BBackbone=Qwen2.5-VL-7B2026.05 | 55.78 | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B2026.05 | 55.37 | — | — | — | — | — | — | — | — | |
| FineViT-VLLanguage Backbone=Qwen3 1.7B LLM, Evaluation Framework=VLMEvalKit [69]2026.03 | 55.05 | — | — | — | — | — | — | — | — | |
| ThinkLite-VL-7BBackbone=Qwen2.5-VL-7B2026.05 | 54.71 | — | — | — | — | — | — | — | — | |
| TroLModel size=7B2024.06 | 54.7 | — | — | — | — | — | — | — | — | |
| Phantom-3.8BModel Size=3.8B2024.09 | 54.4 | — | — | — | — | — | — | — | — |