Multimodal Understanding on SEED Benchmark (test)
73.8Avg Score (All)MetaQuery-L
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| MetaQuery-LParams=3.0B | 3.2B2025.09 | 73.8 | — | — | |
| DIM-4.6B-T2I/EditParams=3.0B | 1.6B2025.09 | 73.8 | — | — | |
| Janus-Pro-7BParams=7.0B2025.09 | 72.1 | — | — | |
| Show-o2-7BParams=7.0B2025.09 | 69.8 | — | — | |
| Emu3-GenParams=8.0B2025.09 | 68.2 | — | — | |
| JanusParams=1.3B2025.09 | 63.7 | — | — | |
| MaVEnVision Encoder=ViT-L + SEED (1.3B), Language Model=Vicuna (7B), zero-shot=true2024.08 | 60.89 | 65.85 | 42.11 | |
| LLaVA-1.5Vision Encoder=ViT-L, Language Model=Vicuna (7B), zero-shot=true2024.08 | 58.6 | 66.1 | 37.3 | |
| InstructBLIPVision Encoder=ViT-g (1.3B), Language Model=Vicuna (7B), zero-shot=true2024.08 | 53.4 | 58.8 | 38.1 | |
| BLIP-2Vision Encoder=ViT-g (1.3B), Language Model=Vicuna (7B), zero-shot=true2024.08 | 46.4 | 49.7 | 36.7 | |
| MiniGPT-4Vision Encoder=ViT-g (1.3B), Language Model=Vicuna (7B), zero-shot=true2024.08 | 42.8 | 47.4 | 29.9 | |
| mPLUG-OwlVision Encoder=ViT-L (0.3B), Language Model=LLaMA (7B), zero-shot=true2024.08 | 34 | 37.9 | 23 | |
| OtterVision Encoder=ViT-L (0.3B), Language Model=LLAMA (7B), zero-shot=true2024.08 | 33.9 | 35.2 | 30.4 | |
| LLaVAVision Encoder=ViT-L (0.3B), Language Model=Vicuna (7B), zero-shot=true2024.08 | 33.5 | 37 | 23.8 | |
| LLaMA-Adapter-v2Vision Encoder=ViT-L (0.3B), Language Model=LLaMA (7B), zero-shot=true2024.08 | 32.7 | 35.2 | 25.8 |