Multimodal Benchmarking on MMBench
83.4ScoreLongVILA-7B (S3)
Evaluation Results
| Method | Links | |
|---|---|---|
| LongVILA-7B (S3)LLM=Qwen2-7B, Resolution=dynamic2024.08 | 83.4 | |
| Meteor2024.05 | 82.9 | |
| Brote-IM-XXL#Param LLM=11B, Setting=Zero-shot, Input=Single-image2024.02 | 77.34 | |
| Brote-EX-XXL#Param LLM=11B, Setting=Zero-shot, Input=Single-image2024.02 | 76.67 | |
| MMICL-XXL#Param LLM=11B, Setting=Zero-shot, Input=Single-image2024.02 | 76.58 | |
| Cambrian-1-8BLLM=Llama 3-8B, Resolution=10242024.08 | 75.9 | |
| INF-LLaVA*Source=Ours, larger_dataset=true2024.07 | 74.38 | |
| Brote-IM-XL#Param LLM=3B, Setting=Zero-shot, Input=Single-image2024.02 | 74.29 | |
| GRACEData %=20%2026.05 | 73.8 | |
| StepmaxData %=20%2026.05 | 73.7 | |
| LongestData %=20%2026.05 | 73.6 | |
| Brote-EX-XL#Param LLM=3B, Setting=Zero-shot, Input=Single-image2024.02 | 73.27 | |
| MMICL-XL#Param LLM=3B, Setting=Zero-shot, Input=Single-image2024.02 | 73.11 | |
| CADCData %=20%2026.05 | 73.1 | |
| Mini-Gemini-HD-8BLLM=Llama 3-8B, Resolution=15362024.08 | 72.7 | |
| RandomData %=20%2026.05 | 72.6 | |
| ICONSData %=20%2026.05 | 71.5 | |
| SPHINX-MoELLM Training #P=8x7B, Input image resolution (Res.)=448, Pre-training samples (PT)=-, Instruction tuning samples (IT)=15.3M2024.08 | 71.3 | |
| GRACEData %=5%2026.05 | 71.1 | |
| GRACEData %=15%2026.05 | 71.1 | |
| INF-LLaVASource=Ours2024.07 | 70.35 | |
| InstructBLIP-XXL#Param LLM=11B, Setting=Zero-shot, Input=Single-image2024.02 | 70.34 | |
| VILALLM=Llama 2-13B, Resolution=3362024.08 | 70.3 | |
| LESSData %=20%2026.05 | 70.2 | |
| HoneybeeSource=CVPR’242024.07 | 70.1 | |
| InstructBLIP-XL#Param LLM=3B, Setting=Zero-shot, Input=Single-image2024.02 | 69.68 | |
| Prompt HighlighterSource=CVPR’242024.07 | 69.5 | |
| GRACEData %=10%2026.05 | 69.5 | |
| FullData %=100%2026.05 | 69.3 | |
| VILALLM=Llama 2-7B, Resolution=3362024.08 | 68.9 | |
| ConvLLaVASource=arXiv’242024.07 | 68.7 | |
| AcFormerSource=arXiv’242024.07 | 68.4 | |
| MoE-LLaVA-2.7Bx4LLM Training #P=5B, Input image resolution (Res.)=384, Pre-training samples (PT)=558K, Instruction tuning samples (IT)=1.6M2024.08 | 68 | |
| MoExtendLLM Training #P=3B, Input image resolution (Res.)=336, Pre-training samples (PT)=558K, Instruction tuning samples (IT)=665K2024.08 | 67.8 | |
| LLaVA-1.5LLM Training #P=13B, Input image resolution (Res.)=336, Pre-training samples (PT)=558K, Instruction tuning samples (IT)=665K2024.08 | 67.7 | |
| LLaVA-1.5LLM=Vicuna-1.5-13B, Resolution=3362024.08 | 67.7 | |
| HyperLLaVALLM Training #P=7B, Input image resolution (Res.)=336, Pre-training samples (PT)=558K, Instruction tuning samples (IT)=665K2024.08 | 65.9 | |
| TIVESource=arXiv’242024.07 | 65.8 | |
| RoboMamba2024.05 | 65.7 | |
| MMICL-XXL (BLIP-2)#Param LLM=11B, Setting=Zero-shot, Input=Single-image2024.02 | 65.24 | |
| MoE-LLaVA-2.7Bx4LLM Training #P=5B, Input image resolution (Res.)=336, Pre-training samples (PT)=558K, Instruction tuning samples (IT)=1.6M2024.08 | 65.2 | |
| GenLLaVASource=arXiv’242024.07 | 65 | |
| LLaVA-NeXT-8BLLM=Llama 3-8B, Resolution=6722024.08 | 64.6 | |
| LLaVA1.5Source=CVPR’242024.07 | 64.3 | |
| LLaVA-1.5LLM Training #P=7B, Input image resolution (Res.)=336, Pre-training samples (PT)=558K, Instruction tuning samples (IT)=665K2024.08 | 64.3 | |
| LLaVA-1.5LLM=Vicuna-1.5-7B, Resolution=3362024.08 | 64.3 | |
| InstructBLIPLLM=V-7B, Res.=2242023.11 | 60.9 | |
| Video-LLaVALLM=V-7B, Res.=2242023.11 | 60.9 | |
| Qwen-VL-7B-ChatLLM Training #P=7B, Input image resolution (Res.)=448, Pre-training samples (PT)=1.4B, Instruction tuning samples (IT)=50M2024.08 | 60.6 | |
| Qwen-VL-ChatLLM=Qwen-7B, Resolution=4482024.08 | 60.6 | |
| OneLLMSource=CVPR’242024.07 | 60 | |
| LLaVA-PhiSource=arXiv’242024.07 | 59.8 | |
| LLaVA-1.5†LLM=V-7B, Res.=224, Encoder=LanguageBind-Image2023.11 | 59.5 | |
| ShikraSource=arXiv’232024.07 | 58.8 | |
| ShikraLLM Training #P=13B, Input image resolution (Res.)=224, Pre-training samples (PT)=600K, Instruction tuning samples (IT)=5.5M2024.08 | 58.8 | |
| OtterHD-8BSource=arXiv’232024.07 | 58.3 | |
| DreamLLMSource=ICLR’242024.07 | 58.2 | |
| InstructBLIPLLM=V-13B, Res.=2242023.11 | 58.2 | |
| VL-Mamba2024.05 | 57 | |
| IDEFICS-80BLLM Training #P=65B, Input image resolution (Res.)=224, Pre-training samples (PT)=353M, Instruction tuning samples (IT)=1M2024.08 | 54.5 | |
| Otter#Param LLM=7B, Setting=Zero-shot, Input=Single-image2024.02 | 48.3 | |
| IDEFICS-9BLLM Training #P=7B, Input image resolution (Res.)=224, Pre-training samples (PT)=353M, Instruction tuning samples (IT)=1M2024.08 | 48.2 | |
| LLaVASource=NeurIPS’232024.07 | 38.7 | |
| GILLSource=NeurIPS’232024.07 | 38.2 | |
| Qwen-VL-7BLLM Training #P=7B, Input image resolution (Res.)=448, Pre-training samples (PT)=1.4B, Instruction tuning samples (IT)=50M2024.08 | 38.2 | |
| Qwen-VLLLM=Qwen-7B, Resolution=4482024.08 | 38.2 | |
| BLIP-2LLM=V-13B, Res.=2242023.11 | 38.1 | |
| LLaVA#Param LLM=7B, Setting=Zero-shot, Input=Single-image2024.02 | 36.2 | |
| AnyGPTSource=arXiv’242024.07 | 36 | |
| InstructBLIP-7BLLM Training #P=7B, Input image resolution (Res.)=224, Pre-training samples (PT)=129M, Instruction tuning samples (IT)=1.2M2024.08 | 36 | |
| InstructBLIPLLM=Vicuna-7B, Resolution=2242024.08 | 36 | |
| MiniGPT-4LLM=L-7B, Res.=2242023.11 | 23 | |
| MGIESource=ICLR’242024.07 | 6.6 |