Visual Reasoning and Instruction Following on MM-Vet
75.2Overall ScoreVLsI-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| VLsI-7BParameters=7B2024.12 | 75.2 | |
| Phantom-7BParameters=7B2024.12 | 70.8 | |
| Qwen2-VL-7BParameters=7B2024.12 | 62 | |
| CogVLM2-8BParameters=8B2024.12 | 60.4 | |
| MiniCPM-V2.6-8BParameters=8B2024.12 | 60 | |
| LLaVA-OneVision-7BParameters=7B2024.12 | 57.5 | |
| TroL-7BParameters=7B2024.12 | 54.7 | |
| InternVL2-8BParameters=8B2024.12 | 54.2 | |
| MiniGemini-HD-13BParameters=13B2024.12 | 50.5 | |
| VILA2-8BParameters=8B2024.12 | 50 | |
| LLaVA-NeXT-13BParameters=13B2024.12 | 47.3 | |
| MM1-MoE-7B×32Parameters=7B×322024.12 | 45.2 | |
| VILA1.5-13BParameters=13B2024.12 | 44.3 | |
| LLaVA-NeXT-7BParameters=7B2024.12 | 43.9 | |
| VILA1.5-8BParameters=8B2024.12 | 43.2 | |
| MM1-7BParameters=7B2024.12 | 42.1 | |
| MiniGemini-HD-7BParameters=7B2024.12 | 41.3 | |
| DPOAlignment method=DPO, Base Model=LLaVA-1.5-13b2024.02 | 41.2 | |
| Rejection-samplingAlignment method=Rejection-sampling, Base Model=LLaVA-1.5-13b2024.02 | 38 | |
| LLAVA-RLHF-13bModel=LLAVA-RLHF-13b2024.02 | 37.2 | |
| Standard SFTAlignment method=Standard SFT, Base Model=LLaVA-1.5-13b2024.02 | 36.5 | |
| LLaVA-1.5-13bModel=LLaVA-1.5-13b2024.02 | 36.3 | |
| SteerLMAlignment method=SteerLM, Base Model=LLaVA-1.5-13b2024.02 | 35.2 |