Visual Question Answering on MMMU (val)
69.1AccuracyGPT-4o
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-4oModel Modality Type=Proprietary2024.10 | 69.1 | |
| GPT-4o-miniModel Modality Type=Proprietary2024.10 | 60 | |
| V3Fusion-RectifyModel ID=2352026.03 | 56.09 | |
| V3Fusion-MLPModel ID=2352026.03 | 55.07 | |
| InternVL2.5-MPOModel Size=8B2025.01 | 52.8 | |
| Qwen2 VLModel Scale=7B, Model Modality Type=Vision-language2024.10 | 52.7 | |
| Qwen2.5-VL-7b-InstructModel ID=52026.03 | 51.55 | |
| Intern-VL2-8bModel ID=62026.03 | 51.3 | |
| V3Fusion-LEDModel ID=2352026.03 | 51.3 | |
| Qwen2-VLModel Size=8B2025.01 | 50.9 | |
| SAIL-VLModel Size=8B2025.01 | 48.2 | |
| DeepSeekVL-2Model Size=8B2025.01 | 47.6 | |
| Baichuan-omniModel Scale=7B, Model Modality Type=Omni-modal2024.10 | 47.3 | |
| MiniCPM-Llama3-V 2.5Model Scale=8B, Model Modality Type=Vision-language2024.10 | 45.8 | |
| VITAModel Scale=8x7B, Model Modality Type=Omni-modal2024.10 | 45.3 | |
| LLaVA-VTLLM=Vicuna-13B, #PT=558K, #IT=665K, Representation=E + T2024.03 | 41.9 | |
| InternVL2.5-MPOModel Size=2B2025.01 | 41.2 | |
| LLaVA-CapLLM=Vicuna-13B, #PT=558K, #IT=665K, Representation=E + T2024.03 | 40.6 | |
| LLaVA-SGLLM=Vicuna-13B, #PT=558K, #IT=665K, Representation=E + T2024.03 | 40.6 | |
| SAIL-VLModel Size=2B2025.01 | 40.1 | |
| Qwen2-VLModel Size=2B2025.01 | 39.9 | |
| DeepSeek-VL2-SmallModel ID=42026.03 | 39.75 | |
| Vicuna-VTLLM=Vicuna-13B, #IT=665K, Representation=Text (T)2024.03 | 39.6 | |
| DeepSeekVL-2Model Size=2B2025.01 | 39.6 | |
| LLaVA-1.5 (V-13B)LLM=Vicuna-13B, #PT=558K, #IT=665K, Representation=Visual Embeddings (E)2024.03 | 39.5 | |
| Vicuna-CapLLM=Vicuna-13B, #IT=665K, Representation=Text (T)2024.03 | 39.1 | |
| DeepSeek-VL2-TinyModel ID=32026.03 | 39.01 | |
| LLaVA-DCapLLM=Vicuna-13B, #PT=558K, #IT=665K, Representation=E + T2024.03 | 38.6 | |
| Vicuna-DCapLLM=Vicuna-13B, #IT=665K, Representation=Text (T)2024.03 | 37.4 | |
| LlaVA-v1.6-Vicuna-13bModel ID=12026.03 | 37.02 | |
| Vicuna-SGLLM=Vicuna-13B, #IT=665K, Representation=Text (T)2024.03 | 36.5 | |
| InstructBLIP (V-7B)LLM=Vicuna-7B, #PT=129M, #IT=1.2M, Representation=Visual Embeddings (E)2024.03 | 36 | |
| Qwen-VL-ChatParameters=7B2024.03 | 35.9 | |
| LlaVA-v1.6-Vicuna-7bModel ID=22026.03 | 35.9 | |
| LLaVA-v1.5Parameters=7B2024.03 | 35.6 | |
| LumenParameters=7B2024.03 | 35.2 | |
| LLaVA-1.5 (V-7B)LLM=Vicuna-7B, #PT=558K, #IT=665K, Representation=Visual Embeddings (E)2024.03 | 33.9 | |
| InstructBLIPParameters=7B2024.03 | 32.9 |