Visual Question Answering on Visual7W (test)
85.33Average AccuracyShikra
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Shikra2023.06 | 85.33 | — | — | — | — | — | — | |
| Our SAT->STBackbone=BERTB, Training Strategy=All-Task Pretraining + Single-Task Finetuning, Number of Parameters=3B, Number of Models=12 x 250M2019.12 | 83.35 | — | — | — | — | — | — | |
| Lu et al.*finetuning=without2023.06 | 83.35 | — | — | — | — | — | — | |
| Our SATBackbone=BERTB, Training Strategy=All-Task (AT), Number of Parameters=270M, Number of Models=1 x 270M2019.12 | 82.75 | — | — | — | — | — | — | |
| Lu et al.2023.06 | 82.75 | — | — | — | — | — | — | |
| SOTA [16]2019.12 | 72.53 | — | — | — | — | — | — | |
| Hu et al.2023.06 | 72.53 | — | — | — | — | — | — | |
| MCB+Att.pooling=Multimodal Compact Bilinear, mechanism=Attention2016.06 | 62.2 | 60.3 | 70.4 | 79.5 | 69.2 | 58.2 | 51.1 | |
| Zhu et al.2023.06 | 56.1 | — | — | — | — | — | — | |
| Zhu et al.type=Baseline2016.06 | 54.3 | 51.5 | 57 | 75 | 59.5 | 55.5 | 49.8 | |
| Concat+Att.pooling=Concatenation, mechanism=Attention2016.06 | 52.8 | 47.8 | 56.9 | 74.1 | 62.3 | 52.7 | 51.2 | |
| Qwen-VL + SC-Tunezero-shot=true2024.03 | 46.09 | — | — | — | — | — | — | |
| MiniGPT-v2 + SC-Tunezero-shot=true2024.03 | 35.29 | — | — | — | — | — | — | |
| Qwen-VLzero-shot=true2024.03 | 34.77 | — | — | — | — | — | — | |
| MiniGPT-v2zero-shot=true2024.03 | 27.53 | — | — | — | — | — | — |