Multimodal Instruction Following on LLaVA-Bench In-the-Wild
93.1ScoreGPT-4V
Evaluation Results
| Method | Links | |
|---|---|---|
| GPT-4V2024.01 | 93.1 | |
| InternLM-XComposer2Number of Parameters=7B2024.01 | 81.8 | |
| Gemini-Pro2024.01 | 79.9 | |
| LyricsLLM Backbone=Vicuna-13B2023.12 | 76.9 | |
| CSRBackbone=LLaVA-1.5-13B2024.05 | 74.7 | |
| CogVLMNumber of Parameters=17B2024.01 | 73.9 | |
| Qwen-VL-Plus2024.01 | 73.7 | |
| ShareGPT4VLLM Backbone=Vicuna-7B2023.12 | 72.6 | |
| CSRBackbone=LLaVA-1.5-7B2024.05 | 71.1 | |
| LLaVA-1.5LLM Backbone=Vicuna-13B2023.12 | 70.7 | |
| LLaVA-1.5-13BModel Size=13B2024.05 | 70.7 | |
| SemDeDupTraining Data Percentage=20%, Data Selection Constraint=Partial Data Used before Selection2026.04 | 70 | |
| POVIDBackbone=LLaVA-1.5-7B2024.05 | 68.7 | |
| PerplexityTraining Data Percentage=20%, Data Selection Constraint=Partial Data Used before Selection2026.04 | 68.3 | |
| FullTraining Data Percentage=100%, Data Selection Constraint=None (Oracle)2026.04 | 67.9 | |
| Qwen-VL-ChatLLM Backbone=Qwen-7B2023.12 | 67.7 | |
| COINCIDETraining Data Percentage=20%, Data Selection Constraint=Full Data Used before Selection2026.04 | 67.3 | |
| CLIP-ScoreTraining Data Percentage=20%, Data Selection Constraint=Partial Data Used before Selection2026.04 | 66.2 | |
| ICONSTraining Data Percentage=20%, Data Selection Constraint=Full Data Used before Selection2026.04 | 66.1 | |
| DOSETraining Data Percentage=20%, Data Selection Constraint=Partial Data Used before Selection2026.04 | 65.8 | |
| Self-rewardingBackbone=LLaVA-1.5-13B2024.05 | 65.6 | |
| RLHF-VBackbone=LLaVA-1.5-7B2024.05 | 65.4 | |
| RandomTraining Data Percentage=20%, Data Selection Constraint=Partial Data Used before Selection2026.04 | 65 | |
| EL2NTraining Data Percentage=20%, Data Selection Constraint=Partial Data Used before Selection2026.04 | 64.9 | |
| Self-FilterTraining Data Percentage=20%, Data Selection Constraint=Partial Data Used before Selection2026.04 | 64.9 | |
| D2-PruningTraining Data Percentage=20%, Data Selection Constraint=Partial Data Used before Selection2026.04 | 63.9 | |
| Human-PreferBackbone=LLaVA-1.5-7B2024.05 | 63.7 | |
| LLaVA-1.5-7BModel Size=7B2024.05 | 63.4 | |
| Self-SupTraining Data Percentage=20%, Data Selection Constraint=Partial Data Used before Selection2026.04 | 63.3 | |
| LLaVALLM Backbone=Vicuna-7B2023.12 | 63 | |
| VlfeedbackBackbone=LLaVA-1.5-7B2024.05 | 62.1 | |
| Self-rewardingBackbone=LLaVA-1.5-7B2024.05 | 61.2 | |
| InstructBLIPLLM Backbone=Vicuna-7B2023.12 | 60.9 | |
| IDEFICS-80BLLM Backbone=LLaMA-65B2023.12 | 56.9 | |
| BLIP-2LLM Backbone=FLAN-T52023.12 | 38.1 |