Multimodal Understanding on SEED-I
77.2AccuracyPVC InternVL2 (Ours)
Evaluation Results
| Method | Links | |
|---|---|---|
| PVC InternVL2 (Ours)Size=8B, #token /image tile=642024.12 | 77.2 | |
| GPT-4o2024.12 | 76.2 | |
| InternVL2Size=8B, #token /image tile=2562024.12 | 76.2 | |
| IXC-2.5Size=7B, #token /image tile=4002024.12 | 75.4 | |
| LLaVA-OVSize=7B, #token /image tile=7292024.12 | 75.4 | |
| PVC InternVL2 (Ours)Size=2B, #token /image tile=642024.12 | 73.2 | |
| InternVL2Size=2B, #token /image tile=2562024.12 | 71.6 | |
| SoM-LLaVA-1.5LLM=Vicuna-13B, Res.=336, Pre-Data=558K, IT-Data=695K2024.04 | 69.6 | |
| SoM-LLaVA-1.5-TLLM=Vicuna-13B, Res.=336, Pre-Data=558K, IT-Data=695K, Tagged Images=true2024.04 | 69.5 | |
| SPHINXLLM=LLAMA2-7B, Res.=2242024.04 | 69.1 | |
| LLaVA-1.5LLM=Vicuna-13B, Res.=336, Pre-Data=558K, IT-Data=665K2024.04 | 68.2 | |
| Qwen-VL-ChatLLM=Qwen-7B, Res.=448, Pre-Data=1.4B+, IT-Data=50M+2024.04 | 65.4 | |
| mPLUG-Owl-2LLM=LLAMA2-7B, Res.=448, Pre-Data=348M2024.04 | 64.1 | |
| Qwen-VLLLM=Qwen-7B, Res.=448, Pre-Data=1.4B+, IT-Data=50M+2024.04 | 62.3 | |
| InstructBLIPLLM=Vicuna-7B, Res.=224, Pre-Data=129M, IT-Data=1.2M2024.04 | 58.8 | |
| InstructBLIPLLM=Vicuna-13B, Res.=224, Pre-Data=129M, IT-Data=1.2M2024.04 | 58.2 | |
| GPT-4V2024.12 | 49.9 | |
| BLIP-2LLM=Vicuna-13B, Res.=224, Pre-Data=129M2024.04 | 49.7 | |
| LLaMA-Adapter-V2LLM=LLAMA2-7B, Res.=3362024.04 | 35.2 |