Multi-modal Understanding on SEED-Bench (overall)
62.9Overall ScoreCSR
Evaluation Results
| Method | Links | |
|---|---|---|
| CSRBackbone=LLaVA-1.5-13B2024.05 | 62.9 | |
| Self-rewardingBackbone=LLaVA-1.5-13B2024.05 | 62.8 | |
| LLaVA-1.5-13BModel Size=13B2024.05 | 61.6 | |
| CogVLM2025.03 | 61.22 | |
| Task Arithmeticunsupervised=false2025.03 | 60.85 | |
| CSRBackbone=LLaVA-1.5-7B2024.05 | 60.3 | |
| POVIDBackbone=LLaVA-1.5-7B2024.05 | 60.2 | |
| AdaMMSunsupervised=true2025.03 | 60.16 | |
| RLHF-VBackbone=LLaVA-1.5-7B2024.05 | 60.1 | |
| Self-rewardingBackbone=LLaVA-1.5-7B2024.05 | 60 | |
| DARE-Linearunsupervised=false2025.03 | 59.81 | |
| mPLUG-Owl2(base)2025.03 | 59.41 | |
| VlfeedbackBackbone=LLaVA-1.5-7B2024.05 | 59.3 | |
| LLaVA-1.5-HACLVision Encoder=VIT-L (0.3B), Language Model=Vicuna (7B), Zero-shot=true2023.12 | 58.9 | |
| LLaVA-1.5Vision Encoder=ViT-L (0.3B), Language Model=Vicuna (7B), Zero-shot=true2023.11 | 58.6 | |
| LLaVA-1.5Vision Encoder=VIT-L (0.3B), Language Model=Vicuna (7B), Zero-shot=true2023.12 | 58.6 | |
| LLaVA-1.5-7BModel Size=7B2024.05 | 58.6 | |
| Qwen-VL-ChatVision Encoder=ViT-G (1.9B), Language Model=Qwen (7B), Zero-shot=true2023.11 | 58.2 | |
| Qwen-VL-ChatVision Encoder=ViT-G (1.9B), Language Model=Qwen (7B), Zero-shot=true2023.12 | 58.2 | |
| Human-PreferBackbone=LLaVA-1.5-7B2024.05 | 58.1 | |
| mPLUG-Owl2Vision Encoder=ViT-L (0.3B), Language Model=LLAMA (7B), Zero-shot=true2023.11 | 57.8 | |
| DARE-Tiesunsupervised=false2025.03 | 57.62 | |
| InstructBLIPVision Encoder=ViT-g (1.3B), Language Model=Vicuna (7B), Zero-shot=true2023.11 | 53.4 | |
| InstructBLIPVision Encoder=ViT-g (1.3B), Language Model=Vicuna (7B), Zero-shot=true2023.12 | 53.4 | |
| Ties-Mergingunsupervised=false2025.03 | 52.32 | |
| MetaGPTunsupervised=true2025.03 | 50.81 | |
| BLIP-2Vision Encoder=ViT-g (1.3B), Language Model=Vicuna (7B), Zero-shot=true2023.11 | 46.4 | |
| BLIP-2Vision Encoder=ViT-g (1.3B), Language Model=Vicuna (7B), Zero-shot=true2023.12 | 46.4 | |
| MiniGPT-4Vision Encoder=ViT-g (1.3B), Language Model=Vicuna (7B), Zero-shot=true2023.11 | 42.8 | |
| MiniGPT-4Vision Encoder=ViT-g (1.3B), Language Model=Vicuna (7B), Zero-shot=true2023.12 | 42.8 | |
| MiniGPT-4-HACLVision Encoder=ViT-g (1.3B), Language Model=Vicuna (7B), Zero-shot=true2023.12 | 42.5 | |
| mPLUG-OwlVision Encoder=ViT-L (0.3B), Language Model=LLaMA (7B), Zero-shot=true2023.11 | 34 | |
| mPLUG-OwlVision Encoder=ViT-L (0.3B), Language Model=LLaMA (7B), Zero-shot=true2023.12 | 34 | |
| LLaVA-HACLVision Encoder=ViT-L (0.3B), Language Model=Vicuna (7B), Zero-shot=true2023.12 | 33.9 | |
| LLaVAVision Encoder=ViT-L (0.3B), Language Model=Vicuna (7B), Zero-shot=true2023.11 | 33.5 | |
| LLaVAVision Encoder=VIT-L (0.3B), Language Model=Vicuna (7B), Zero-shot=true2023.12 | 33.5 | |
| OtterVision Encoder=ViT-L (0.3B), Language Model=LLaMA (7B), Zero-shot=true2023.11 | 32.9 | |
| OtterVision Encoder=VIT-L (0.3B), Language Model=LLaMA (7B), Zero-shot=true2023.12 | 32.9 | |
| LLaMA-Adapter-v2Vision Encoder=ViT-L (0.3B), Language Model=LLaMA (7B), Zero-shot=true2023.11 | 32.7 | |
| LLaMA-Adapter-v2Vision Encoder=ViT-L (0.3B), Language Model=LLaMA (7B), Zero-shot=true2023.12 | 32.7 |