Multimodal Question Answering on MM-Vet
68.3Total ScoreQwen3-VL-4B
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Qwen3-VL-4BModel Scale=4B2026.02 | 68.3 | — | — | — | — | — | — | |
| DenseMLLM-4BModel Scale=4B2026.02 | 64.6 | — | — | — | — | — | — | |
| LLaVANextData=OmniAlign-Vmix, LLM=Qwen2.5-32B2025.02 | 56.9 | — | — | — | — | — | — | |
| LLaVANextData=OmniAlign-Vmix, LLM=InternLM2.5-7B2025.02 | 47.7 | — | — | — | — | — | — | |
| LLaVANextData=LLaVANext-778k, LLM=Qwen2.5-32B2025.02 | 47.7 | — | — | — | — | — | — | |
| LLaVAData=OmniAlign-Vmix, LLM=InternLM2.5-7B2025.02 | 43.5 | — | — | — | — | — | — | |
| LLaVANextData=LLaVANext-778k, LLM=InternLM2.5-7B2025.02 | 41.8 | — | — | — | — | — | — | |
| LLaVAData=LLaVANext-778k, LLM=InternLM2.5-7B2025.02 | 41.2 | — | — | — | — | — | — | |
| LLaVA-VTLLM=Vicuna-13B, #PT=558K, #IT=665K, Representation=E + T2024.03 | 39.8 | — | — | — | — | — | — | |
| LLaVA-COCO-13BMulti-round=Yes2024.01 | 37.5 | 41.2 | 27.9 | 27.4 | 25.4 | 31.1 | 15 | |
| LLaVA-DCapLLM=Vicuna-13B, #PT=558K, #IT=665K, Representation=E + T2024.03 | 36.7 | — | — | — | — | — | — | |
| LLaVA-CapLLM=Vicuna-13B, #PT=558K, #IT=665K, Representation=E + T2024.03 | 36.3 | — | — | — | — | — | — | |
| LLaVA-SGLLM=Vicuna-13B, #PT=558K, #IT=665K, Representation=E + T2024.03 | 36.1 | — | — | — | — | — | — | |
| LLaVA-1.5-13BMulti-round=No2024.01 | 36 | 40.3 | 28.3 | 22.6 | 23.9 | 34.9 | 7.7 | |
| LLaVA-COCO-13BMulti-round=No2024.01 | 35.4 | 39.3 | 28.5 | 24.6 | 24.4 | 29.1 | 11.2 | |
| LLaVA-1.5 (V-13B)LLM=Vicuna-13B, #PT=558K, #IT=665K, Representation=Visual Embeddings (E)2024.03 | 35.4 | — | — | — | — | — | — | |
| Vicuna-VTLLM=Vicuna-13B, #IT=665K, Representation=Text (T)2024.03 | 30.7 | — | — | — | — | — | — | |
| LLaVA-1.5 (V-7B)LLM=Vicuna-7B, #PT=558K, #IT=665K, Representation=Visual Embeddings (E)2024.03 | 30.5 | — | — | — | — | — | — | |
| LLaVA-1.5-13BMulti-round=Yes2024.01 | 29.2 | 31.3 | 25.7 | 7.5 | 9.3 | 33.3 | 11.2 | |
| Vicuna-SGLLM=Vicuna-13B, #IT=665K, Representation=Text (T)2024.03 | 28.1 | — | — | — | — | — | — | |
| Vicuna-DCapLLM=Vicuna-13B, #IT=665K, Representation=Text (T)2024.03 | 27.1 | — | — | — | — | — | — | |
| InstructBLIP (V-7B)LLM=Vicuna-7B, #PT=129M, #IT=1.2M, Representation=Visual Embeddings (E)2024.03 | 26.2 | — | — | — | — | — | — | |
| InstructBLIP (V-13B)LLM=Vicuna-13B, #PT=129M, #IT=1.2M, Representation=Visual Embeddings (E)2024.03 | 25.6 | — | — | — | — | — | — | |
| Vicuna-CapLLM=Vicuna-13B, #IT=665K, Representation=Text (T)2024.03 | 23 | — | — | — | — | — | — |