General VQA on MMStar
80.1AccuracyQwen3.5-Plus
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3.5-PlusParameter Count=397B2026.06 | 80.1 | |
| K2.5Parameter Count=≈1T2026.06 | 79.8 | |
| Qwen3-VL-32BDecoding Mode=Thinking, Model Scale=32B2026.06 | 77.1 | |
| Qwen3-VL-32BDecoding Mode=Instruct, Model Scale=32B2026.06 | 76 | |
| Qwen3-VL-8BDecoding Mode=Thinking, Model Scale=8B2026.06 | 75.2 | |
| Opus-4.6Parameter Count=≈800B2026.06 | 74.5 | |
| Jigsaw-R1-8BModel Category=Template-based Self-evolution Methods, Scale=8B2026.04 | 74.3 | |
| EVE (Ours-8B-iter4)Model Category=Pseudo-label-based Self-evolution Methods, Scale=8B, Iteration=42026.04 | 73.9 | |
| VisPlay-8B-iter3Model Category=Pseudo-label-based Self-evolution Methods, Scale=8B, Iteration=32026.04 | 73.1 | |
| GPT-5.4Parameter Count=≈1.5T2026.06 | 72.7 | |
| Qwen3-VL-8B-InstructModel Category=Open-Source MLLMs, Scale=8B2026.04 | 72.1 | |
| Yuvion VL-32BDecoding Mode=Reasoning, Model Scale=32B2026.06 | 71.3 | |
| Yuvion VL-32BDecoding Mode=Instruct, Model Scale=32B2026.06 | 71 | |
| Qwen2.5-VL-7B AIFmethod=Adaptive Information Flow2026.04 | 70.9 | |
| Qwen2.5-VL-72B2026.04 | 70.8 | |
| MM-Zero-8B-iter3Model Category=Pseudo-label-based Self-evolution Methods, Scale=8B, Iteration=32026.04 | 70.7 | |
| Yuvion VL-8BDecoding Mode=Reasoning, Model Scale=8B2026.06 | 70.5 | |
| Qwen3-VL-8BDecoding Mode=Instruct, Model Scale=8B2026.06 | 70.1 | |
| Qwen3-VLModel Size=4B, Training Stage=instruct2026.05 | 69.8 | |
| M2NoteBackbone=Qwen3-VL-Plus, Evolving Mode=Self-Evolving, Mem=28, Len=4622026.07 | 69.4 | |
| GPT-5 nano (high)Model Category=Closed-Source MLLMs, Scale=high2026.04 | 68.6 | |
| VanillaBackbone=Qwen3-VL-Plus, Evolving Mode=Self-Evolving2026.07 | 68.3 | |
| M2NoteBackbone=GPT-5.4, Evolving Mode=Self-Evolving, Mem=28, Len=4622026.07 | 67.8 | |
| Qwen3-VL-SegModel Size=4B, Training Stage=S-22026.05 | 67.7 | |
| Qwen3-VLModel Size=4B, Training Stage=S-12026.05 | 67.5 | |
| VanillaBackbone=GPT-5.4, Evolving Mode=Self-Evolving2026.07 | 66.5 | |
| InternVL3-9BModel Category=Open-Source MLLMs, Scale=9B2026.04 | 66.3 | |
| Qwen-ViPER-7BModel Category=Pseudo-label-based Self-evolution Methods, Scale=7B2026.04 | 66.2 | |
| LLaVA-OneVision-72BModel Category=Open-Source MLLMs, Scale=72B2026.04 | 65.8 | |
| Yuvion VL-8BDecoding Mode=Instruct, Model Scale=8B2026.06 | 65.6 | |
| Vision-ZeroEvolving Mode=Self-Evolving2026.07 | 65.2 | |
| Claude-3.5 Sonnet2026.04 | 65.1 | |
| VisPlayEvolving Mode=Self-Evolving2026.07 | 65.1 | |
| InternVL-3.5Model Size=4B2026.05 | 65 | |
| CoT+M2NoteBackbone=Qwen3-VL-8B-Instruct, Evolving Mode=Self-Evolving, Mem=32, Len=3152026.07 | 64.9 | |
| GPT-4o2026.04 | 64.7 | |
| GPT-4o-20240513Model Category=Closed-Source MLLMs2026.04 | 64.7 | |
| M2NoteBackbone (Tuning)=Qwen3-VL-8B-Instruct, Backbone (Tuner)=Qwen3-VL-Plus, Evolving Mode=Cross-Model Evolving, Mem=43, Len=4452026.07 | 64 | |
| Qwen2.5-VL-7B2026.04 | 63.9 | |
| M2NoteBackbone=Qwen3-VL-8B-Instruct, Evolving Mode=Self-Evolving, Mem=50, Len=2612026.07 | 63.9 | |
| CoTBackbone=Qwen3-VL-8B-Instruct, Evolving Mode=Self-Evolving2026.07 | 62.9 | |
| VanillaBackbone=Qwen3-VL-8B-Instruct, Evolving Mode=Self-Evolving2026.07 | 62.1 | |
| DPEBackbone=Qwen3-VL-8B-Instruct, Evolving Mode=Self-Evolving2026.07 | 62.1 | |
| VanillaBackbone (Tuning)=Qwen3-VL-8B-Instruct, Backbone (Tuner)=Qwen3-VL-Plus, Evolving Mode=Cross-Model Evolving2026.07 | 62.1 | |
| GPT-5 mini (minimal)Model Category=Closed-Source MLLMs, Scale=minimal2026.04 | 61.3 | |
| FineViT-VLLanguage Backbone=Qwen3 1.7B LLM, Evaluation Framework=VLMEvalKit [69]2026.03 | 60.87 | |
| Intern3.5-VLLanguage Backbone=Qwen3 1.7B LLM, Evaluation Framework=VLMEvalKit [69]2026.03 | 57.33 | |
| Qwen3-VLLanguage Backbone=Qwen3 1.7B LLM, Evaluation Framework=VLMEvalKit [69]2026.03 | 56.6 | |
| Aquila-VLLanguage Backbone=Qwen2.5 1.5B LLM, Evaluation Framework=VLMEvalKit [69]2026.03 | 54.53 | |
| LLaVA-1.5-7B AIFmethod=Adaptive Information Flow2026.04 | 39.5 | |
| LLaVA-1.5-7B2026.04 | 33.1 |