Visual Question Answering on MemeVQA 1.0 (test)
87AccuracyARSENAL
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| ARSENALType=Multimodal, Rationale type=Entity-Specific2024.05 | 87 | 0.58 | 0.17 | 0.53 | 0.56 | 0.48 | 0.932 | |
| ARSENALType=Multimodal, Rationale type=Generic2024.05 | 87 | 0.63 | 0.19 | 0.55 | 0.56 | 0.46 | 0.934 | |
| MM-CoTType=Multimodal, Lecture=Included2024.05 | 69 | 0.59 | 0.13 | 0.54 | 0.51 | 0.49 | 0.895 | |
| MM-CoTType=Multimodal, OCR=Yes2024.05 | 67 | 0.59 | 0.12 | 0.54 | 0.51 | 0.49 | 0.894 | |
| MM-CoTType=Multimodal, Prompt=QCML to A, Rationales=LLaVA-generated2024.05 | 66 | 0.59 | 0.12 | 0.54 | 0.51 | 0.49 | 0.896 | |
| MM-CoTType=Multimodal, OCR=No2024.05 | 59 | 0.58 | 0.13 | 0.53 | 0.5 | 0.47 | 0.891 | |
| T5Type=Unimodal (Text)2024.05 | 53 | 0.59 | 0.15 | 0.44 | 0.41 | 0.35 | 0.901 | |
| ViT + BERTType=Unimodal (Image), Visual Backbone=ViT, Text Backbone=BERT2024.05 | 46 | 0.51 | 0.1 | 0.45 | 0.44 | 0.38 | 0.911 | |
| MM ViT + BERTType=Multimodal, Visual Backbone=ViT, Text Backbone=BERT2024.05 | 45 | 0.51 | 0.11 | 0.46 | 0.44 | 0.38 | 0.911 | |
| MM BEiT + BERTType=Multimodal, Visual Backbone=BEiT, Text Backbone=BERT2024.05 | 44 | 0.48 | 0.09 | 0.45 | 0.45 | 0.39 | 0.91 | |
| ViLTType=Multimodal2024.05 | 43 | — | — | — | — | — | — | |
| BEiT + BERTType=Unimodal (Image), Visual Backbone=BEiT, Text Backbone=BERT2024.05 | 40 | 0.5 | 0.11 | 0.44 | 0.44 | 0.38 | 0.909 | |
| miniGPT4Type=Multimodal, Mode=Zero-shot2024.05 | 32 | 0.09 | 0 | 0.14 | 0.21 | 0.23 | 0.753 | |
| GPT-3.5Type=Unimodal (Text)2024.05 | 28 | — | — | — | — | — | — | |
| miniGPT4Type=Multimodal, Mode=Fine-tuned2024.05 | 28 | 0.12 | 0 | 0.16 | 0.23 | 0.26 | 0.771 | |
| LLaVAType=Multimodal, Mode=Zero-shot2024.05 | — | 0.05 | 0 | 0.09 | 0.17 | 0.18 | 0.837 |