Multimodal Reasoning on LLaVA-Bench Wild
74.5GPT-4 ScoreLLaVA-1.5 + MMInstruct
Evaluation Results
| Method | Links | |
|---|---|---|
| LLaVA-1.5 + MMInstructLLM=Vicuna-13B2024.07 | 74.5 | |
| LLaVA-1.5 + MMInstructLLM=Vicuna-7B2024.07 | 71 | |
| LLaVA-1.5LLM=Vicuna-13B2024.07 | 70.7 | |
| LLaVA-1.5LLM=Vicuna-7B2024.07 | 63.4 | |
| InstructBLIPLLM=Vicuna-7B2024.07 | 60.9 | |
| InstructBLIPLLM=Vicuna-13B2024.07 | 58.2 | |
| BLIP-2LLM=Vicuna-13B2024.07 | 38.1 | |
| LLaVA-VTLLM=V-13B, #PT=558K, #IT=665K, Representation=E + T2024.03 | 0.891 | |
| LLaVA-CapLLM=V-13B, #PT=558K, #IT=665K, Representation=E + T2024.03 | 0.889 | |
| LLaVA-SGLLM=V-13B, #PT=558K, #IT=665K, Representation=E + T2024.03 | 0.868 | |
| LLaVA-1.5LLM=V-13B, #PT=558K, #IT=665K, Representation=Visual Embeddings (E)2024.03 | 0.867 | |
| LLaVA-DCapLLM=V-13B, #PT=558K, #IT=665K, Representation=E + T2024.03 | 0.864 | |
| LLaVA-VTLLM=V-7B, #PT=558K, #IT=665K, Representation=E + T2024.03 | 0.85 | |
| Vicuna-VTLLM=V-13B, #IT=665K, Representation=Text-based (T)2024.03 | 0.825 | |
| Vicuna-VTLLM=V-7B, #IT=665K, Representation=Text-based (T)2024.03 | 0.82 | |
| LLaVA-1.5LLM=V-7B, #PT=558K, #IT=665K, Representation=Visual Embeddings (E)2024.03 | 0.819 | |
| Vicuna-CapLLM=V-13B, #IT=665K, Representation=Text-based (T)2024.03 | 0.792 | |
| Vicuna-DCapLLM=V-13B, #IT=665K, Representation=Text-based (T)2024.03 | 0.776 | |
| Vicuna-SGLLM=V-13B, #IT=665K, Representation=Text-based (T)2024.03 | 0.77 |