Visual Question Answering on ScienceQA Image (test)
96.2AccuracyMutimodal-T-SciQLarge
Evaluation Results
| Method | Links | |
|---|---|---|
| Mutimodal-T-SciQLargeEvaluation Protocol=Supervised Fine-tuning, Trainable params=738M2024.08 | 96.2 | |
| MC-CoT-F-LargeEvaluation Protocol=Supervised Fine-tuning, Trainable params=738M2024.08 | 94.9 | |
| CROME-Vicuna-7BEvaluation Protocol=Supervised Fine-tuning, Trainable params=5.24M2024.08 | 93.2 | |
| PILL-7BEvaluation Protocol=Supervised Fine-tuning, Trainable params=45 M2024.08 | 91.2 | |
| LaVIN-13BEvaluation Protocol=Supervised Fine-tuning, Trainable params=5.4M2024.08 | 90.8 | |
| Prophet++Backbone=mPLUG2023.03 | 90.5 | |
| LaVIN-7BEvaluation Protocol=Supervised Fine-tuning, Trainable params=3.8M2024.08 | 89.4 | |
| ProphetBackbone=mPLUG2023.03 | 88.2 | |
| LLaVA2023.03 | 88 | |
| Human Average2023.03 | 87.5 | |
| LLaMA-AdapterEvaluation Protocol=Supervised Fine-tuning, Trainable params=1.8M2024.08 | 85.2 | |
| MM-COT2023.03 | 82.9 | |
| Dense ConnectorPT+IT=0.5M+0.6M, Res.=384, LLM=Llama3-70B LORA2024.05 | 82.4 | |
| Dense ConnectorPT+IT=0.5M+0.6M, Res.=384, LLM=Yi-34B LORA2024.05 | 80.5 | |
| LLaMA-Adapter2023.03 | 80.3 | |
| InstructBLIP2023.03 | 79.5 | |
| Dense ConnectorPT+IT=1.2M+1.5M, Res.=384 AnyRes, LLM=Yi-34B2024.05 | 78 | |
| Dense ConnectorPT+IT=1.2M+1.5M, Res.=384, LLM=Vicuna-13B2024.05 | 77.1 | |
| mPLUG2023.03 | 77 | |
| Dense ConnectorPT+IT=0.5M+0.6M, Res.=384, LLM=Llama3-8B2024.05 | 75.2 | |
| Dense ConnectorPT+IT=1.2M+1.5M, Res.=384 AnyRes, LLM=Vicuna-13B2024.05 | 75.2 | |
| MobileVLM V2PT+IT=1.2M+3.6M, Res.=336, LLM=Vicuna-7B2024.05 | 74.8 | |
| SPHINX-PlusPT+IT=16M, Res.=448, LLM=Llama2-13B2024.05 | 74.2 | |
| VILAPT+IT=50M+1M, Res.=336, LLM=Llama-2-13B2024.05 | 73.7 | |
| LLaVA-NeXTPT+IT=0.5M+0.7M, Res.=336 AnyRes, LLM=Vicuna-13B2024.05 | 73.6 | |
| LLaVA-LLaMA3PT+IT=0.5M+0.6M, Res.=336, LLM=Llama3-8B2024.05 | 73.3 | |
| Dense ConnectorPT+IT=0.5M+0.6M, Res.=384, LLM=Vicuna-13B2024.05 | 73 | |
| MM1PT+IT=3B+1.4M, Res.=1344, LLM=MM1-7B2024.05 | 72.6 | |
| Mini-GeminiPT+IT=1.2M+1.5M, Res.=336+768, LLM=Vicuna-13B2024.05 | 72.6 | |
| Dense ConnectorPT+IT=1.2M+1.5M, Res.=384 AnyRes, LLM=Vicuna-7B2024.05 | 72 | |
| CuMoPT+IT=0.5M+0.6M, Res.=336, LLM=Mistral-7B2024.05 | 71.7 | |
| LLaVA-v1.5PT+IT=0.5M+0.6M, Res.=336, LLM=Vicuna-13B2024.05 | 71.6 | |
| ShareGPT4VPT+IT=1.2M+0.7M, Res.=336, LLM=Vicuna-13B2024.05 | 71.2 | |
| Dense ConnectorPT+IT=0.5M+0.6M, Res.=384, LLM=Vicuna-7B2024.05 | 70.5 | |
| Dense ConnectorPT+IT=0.5M+0.6M, Res.=384, LLM=Phi2-2.7B2024.05 | 70.3 | |
| LLaVA 1.6-7BEvaluation Protocol=Zero-Shot Performance, Trainable params=7B2024.08 | 70.1 | |
| MobileVLM V2PT+IT=1.2M+3.6M, Res.=336, LLM=ML-2.7B2024.05 | 70 | |
| LLaMA-VIDPT+IT=0.8M+0.7M, Res.=336, LLM=Vicuna-7B2024.05 | 70 | |
| TinyLLaVAPT+IT=0.5M+0.6M, Res.=384, LLM=Phi2-2.7B2024.05 | 69.9 | |
| mPLUG-Owl2PT+IT=348M+1.2M, Res.=448, LLM=Llama2-7B2024.05 | 68.7 | |
| Qwen-VL-ChatPT+IT=1.4B+50M, Res.=448, LLM=Qwen-7B2024.05 | 68.2 | |
| CROME-Vicuna-7BEvaluation Protocol=Zero-Shot Performance, Trainable params=199.85M2024.08 | 61.2 | |
| InstructBLIP-Vicuna-7BEvaluation Protocol=Zero-Shot Performance, Trainable params=188M2024.08 | 60.5 | |
| BLIVA-Vicuna-7BEvaluation Protocol=Zero-Shot Performance, Trainable params=194.61M2024.08 | 57.3 | |
| MCAN2023.03 | 51.2 |