Multi-choice Visual Question Answering on A-OKVQA
82.71AccuracyCVLM
Evaluation Results
| Method | Links | |
|---|---|---|
| CVLMLLM=Qwen-VL2024.02 | 82.71 | |
| BaselineModel=LLaVA-13B, TFLOPS=5.81, FLOPs Ratio=100%2025.03 | 82 | |
| HiMAPModel=LLaVA-13B, TFLOPS=1.36, FLOPs Ratio=23%2025.03 | 81.4 | |
| FastVModel=LLaVA-13B, TFLOPS=3.09, FLOPs Ratio=53%2025.03 | 81.3 | |
| HiMAPModel=InternVL-7B, TFLOPS=0.56, FLOPs Ratio=20%2025.03 | 80.1 | |
| BaselineModel=InternVL-7B, TFLOPS=2.71, FLOPs Ratio=100%2025.03 | 79.6 | |
| CVLM (3M IKPairs, Objects=3)LLM=Vicuna-7B, Objects=32024.02 | 79.21 | |
| FastVModel=InternVL-7B, TFLOPS=1.39, FLOPs Ratio=52%2025.03 | 79.1 | |
| CVLM (3M IKPairs, Objects=5)LLM=Vicuna-7B, Objects=52024.02 | 78.95 | |
| CVLM (3M IKPairs, Objects=8)LLM=Vicuna-7B, Objects=82024.02 | 78.95 | |
| CVLM (3M IKPairs, Objects=1)LLM=Vicuna-7B, Objects=12024.02 | 78.69 | |
| InstructBLIPLLM=Flan-T5-XXL2024.02 | 78.34 | |
| CVLM (3M IKPairs) w/o FKALLM=Vicuna-7B2024.02 | 77.9 | |
| CVLMLLM=Vicuna-7B2024.02 | 77.64 | |
| HiMAPModel=LLaVA-7B, TFLOPS=0.73, FLOPs Ratio=24%2025.03 | 77.2 | |
| CVLM w/o FKALLM=Vicuna-7B2024.02 | 77.12 | |
| FastVModel=LLaVA-7B, TFLOPS=1.56, FLOPs Ratio=54%2025.03 | 77 | |
| InstructBLIPLLM=Flan-T5-XL2024.02 | 76.68 | |
| BaselineModel=LLaVA-7B, TFLOPS=2.98, FLOPs Ratio=100%2025.03 | 76.6 | |
| HiMAPModel=QwenVL-7B, TFLOPS=0.89, FLOPs Ratio=25%2025.03 | 75.9 | |
| BaselineModel=QwenVL-7B, TFLOPS=3.6, FLOPs Ratio=100%2025.03 | 75.7 | |
| FastVModel=QwenVL-7B, TFLOPS=1.9, FLOPs Ratio=53%2025.03 | 75.3 | |
| CVLM w/o (FKA & VKA)LLM=Vicuna-7B2024.02 | 73.97 | |
| LLaVA-v1.5†LLM=Vicuna-7B2024.02 | 73.45 | |
| MC-CoT_BaseVision Model=DETR [11], Text Model=UnifiedQABase [24], Parameters=223M2023.11 | 71 | |
| MINDbaseLearning=Fine-tune2025.12 | 70.6 | |
| BLIP-2Vision Model=CLIP-VIT-LARGE [41], Text Model=FlanT5XXL [15], Parameters=11B2023.11 | 70.2 | |
| Qwen-VLLLM=Qwen2024.02 | 70.04 | |
| GPV-2Vision Model=VinVL [59], Text Model=T5-Base [42], Parameters=300M2023.11 | 60.3 | |
| GPV-2Learning=Fine-tune2025.12 | 60.3 | |
| IPVR (GPT-3)Learning=Few-shot2025.12 | 58.7 | |
| ClipCapLearning=Few-shot2025.12 | 56.9 | |
| KRISPVision Model=Faster R-CNN [44], Text Model=BERT [23], Parameters=200M2023.11 | 51.9 | |
| KRISPLearning=Fine-tune2025.12 | 51.9 | |
| LXMERTVision Model=Transformer [49], Text Model=Transformer [49], Parameters=220M2023.11 | 51.4 | |
| LXMERTLearning=Fine-tune2025.12 | 51.4 | |
| Multimodal-CoT_BaseVision Model=DETR [11], Text Model=UnifiedQABase [24], Parameters=223M2023.11 | 50.6 | |
| Multimodal-CoTbaseLearning=Fine-tune2025.12 | 50.6 | |
| VILBERTVision Model=Faster R-CNN [44], Text Model=BERT [23], Parameters=300M2023.11 | 49.1 | |
| ViLBERTLearning=Fine-tune2025.12 | 49.1 | |
| PythiaVision Model=ResNet [19], Text Model=BERT [23], Parameters=70M2023.11 | 49 | |
| PythiaLearning=Fine-tune2025.12 | 49 | |
| IPVRVision Model=Faster-RCNN [44], Text Model=OPT [62], Parameters=66B2023.11 | 48.6 | |
| IPVR (OPT-66B)Learning=Few-shot2025.12 | 48.6 | |
| PaLM-COTText Model=PaLM [13], Parameters=540B2023.11 | 48.1 | |
| CoTLearning=Few-shot2025.12 | 48.1 | |
| PICaVision Model=VinVL [59], Text Model=GPT-3 [10], Parameters=175B2023.11 | 46.1 | |
| PicaLearning=Few-shot2025.12 | 46.1 | |
| InstructBLIPLLM=Vicuna-7B2024.02 | 45.07 |