Large Vision-Language Model Evaluation on SEED
72.5Overall ScoreCogVLM-Chat
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CogVLM-ChatLLM=Vicuna-7B2023.11 | 72.5 | — | |
| SPHINX-2kLLM=LLaMA2 13B2023.11 | 71.6 | — | |
| Unified-IO2LLM=UIO-2XXL2023.11 | 71.5 | — | |
| mPLUG-Owl2LLM=LLaMA2-7B2023.11 | 64.1 | — | |
| Emu2-ChatLLM=LLaMA-33B2023.11 | 62.8 | — | |
| Qwen-VL-ChatLLM=Qwen-7B2023.11 | 61.8 | — | |
| LLaVA-1.5LLM=Vicuna-13B2023.11 | 61.6 | — | |
| InstructBLIPLLM=Vicuna-7B2023.11 | 58.8 | — | |
| LLaVA-1.5LLM=Vicuna-7B2023.11 | 58.6 | — | |
| IDEFICS-InstructLLM=LLaMA-65B2023.11 | 53.2 | — | |
| DreamLLMLLM=Vicuna-7B2023.11 | 49.9 | — | |
| MiniGPT-4LLM=Vicuna-7B2023.11 | 47.4 | — | |
| OpenFlamingoLLM=MPT-7B2023.11 | 42.7 | — | |
| BLIP-2Param.=14.2B, Res.=224, Data=129M2024.03 | — | 46.4 | |
| InstructBLIPParam.=8.2B, Res.=224, Data=130M, Inference Speed=22.6 t/s2024.03 | — | 53.4 | |
| LLaVA-1.5Param.=7.2B, Res.=336, Data=1.2M, Inference Speed=23.8 t/s2024.03 | — | 58.6 | |
| LLaVA-1.5Param.=13.2B, Res.=336, Data=1.2M2024.03 | — | 61.6 | |
| LLaVA-HRParam.=7.4B, Res.=1024, Data=1.2M, Inference Speed=19.7 t/s2024.03 | — | 64.2 | |
| LLaVA-HRParam.=13.4B, Res.=1024, Data=1.2M, Inference Speed=15.0 t/s2024.03 | — | 64.5 | |
| LLaVA-HR-XParam.=14B, Res.=1024, Data=1.2M, Inference Speed=12.9 t/s2024.03 | — | 65.3 | |
| mPLUG-Owl2Param.=8.2B, Res.=448, Data=400M, Inference Speed=19.6 t/s2024.03 | — | 57.8 | |
| QWen-VL-ChatParam.=9.6B, Res.=448, Data=1.4B, Inference Speed=17.0 t/s2024.03 | — | 58.2 |