Direct-answer Visual Question Answering on A-OKVQA
68.7AccuracyMC-CoT_Base
Evaluation Results
| Method | Links | |
|---|---|---|
| MC-CoT_BaseVision Model=DETR [11], Text Model=UnifiedQABase [24], Parameters=223M2023.11 | 68.7 | |
| LASERModel=Qwen-VL, Decoding=LASER2026.02 | 62.82 | |
| VCDModel=Qwen-VL, Decoding=VCD2026.02 | 61.34 | |
| ViCropModel=Qwen-VL, Decoding=ViCrop2026.02 | 60.12 | |
| SampleModel=Qwen-VL, Decoding=Sample2026.02 | 59.64 | |
| BLIP-2Vision Model=CLIP-VIT-LARGE [41], Text Model=FlanT5XXL [15], Parameters=11B2023.11 | 53.2 | |
| GPV-2Vision Model=VinVL [59], Text Model=T5-Base [42], Parameters=300M2023.11 | 48.6 | |
| IPVRVision Model=Faster-RCNN [44], Text Model=OPT [62], Parameters=66B2023.11 | 46.4 | |
| PICaVision Model=VinVL [59], Text Model=GPT-3 [10], Parameters=175B2023.11 | 42.4 | |
| PaLM-COTText Model=PaLM [13], Parameters=540B2023.11 | 41.5 | |
| KRISPVision Model=Faster R-CNN [44], Text Model=BERT [23], Parameters=200M2023.11 | 33.7 | |
| LXMERTVision Model=Transformer [49], Text Model=Transformer [49], Parameters=220M2023.11 | 30.7 | |
| VILBERTVision Model=Faster R-CNN [44], Text Model=BERT [23], Parameters=300M2023.11 | 30.6 | |
| LASERModel=LLaVA-1.5, Decoding=LASER2026.02 | 28.18 | |
| VCDModel=LLaVA-1.5, Decoding=VCD2026.02 | 25.54 | |
| PythiaVision Model=ResNet [19], Text Model=BERT [23], Parameters=70M2023.11 | 25.2 | |
| ViCropModel=LLaVA-1.5, Decoding=ViCrop2026.02 | 23.95 | |
| SampleModel=LLaVA-1.5, Decoding=Sample2026.02 | 23.63 |