Visual Question Answering on GQA (test-std)
65.65AccuracyVinVL
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| VinVLVisual Encoder=VinVL-ResX152, V&L Pretrain Data=8.9M, V&L Pretrain Epoch=1162021.07 | 65.65 | — | — | |
| OSCAR_VinVLVisual features=VinVL [45], Additional supervision=None, Training images (M)=≈ 5.7, Training sentences (M)=~ 92021.06 | 64.7 | 82.3 | 48.8 | |
| OSCAR+ w/ VinVLModel size=Base2021.01 | 64.65 | — | — | |
| VinVLPre-training img data=VG, COCO, Objects365, SBU Flickr30k, CC, VQA, OpenImagesV5 (5.65M)2021.04 | 64.65 | — | — | |
| NSMVisual features=SG [16], Additional supervision=Scene graph, Training images (M)=≈ 0.1, Training sentences (M)=≈ 12021.06 | 63.2 | 78.9 | 49.3 | |
| NSM2020.04 | 63.17 | — | — | |
| NSM2020.04 | 63.17 | — | — | |
| NSM2020.04 | 63.17 | — | — | |
| NSM2021.01 | 63.17 | — | — | |
| NSM2021.04 | 63.17 | — | — | |
| OursVisual features=VinVL [45], Additional supervision=Program, Training images (M)=≈ 0.1, Training sentences (M)=≈ 152021.06 | 63 | 80.1 | 48 | |
| CLIP-ViLpVisual Encoder=CLIP-Res50x4, V&L Pretrain Data=9.2M, V&L Pretrain Epoch=202021.07 | 62.93 | — | — | |
| MDETRBackbone=EfficientNet-B5, Pre-training img data=VG, COCO, Flickr30k (200k)2021.04 | 62.45 | — | — | |
| Partially⇑Architecture=CLIP-ViL2023.08 | 62.28 | — | — | |
| LXMERT+CATT↑GPU Hours=1536 (1080Ti), 1056 (V100), Pre-training Data (Image / Text)=0.18M / 9.18M2021.03 | 62.07 | — | — | |
| MDETRBackbone=ResNet-101, Pre-training img data=VG, COCO, Flickr30k (200k)2021.04 | 61.99 | — | — | |
| OSCARModel Size=Large2020.04 | 61.62 | — | — | |
| OSCARModel Size=L2020.04 | 61.62 | — | — | |
| OSCARModel Size=Large2020.04 | 61.62 | — | — | |
| OSCARModel size=Base2021.01 | 61.62 | — | — | |
| OSCARPre-training img data=VG, COCO, Flickr, SBU (4.3M)2021.04 | 61.62 | — | — | |
| Oscar BasePretraining Images=4M2021.02 | 61.6 | — | — | |
| Fully-FTArchitecture=CLIP-ViL2023.08 | 61.44 | — | — | |
| Fully-FTParams. (M)=236.8, Mem. (G)=82.02026.04 | 61.44 | — | — | |
| Fully-FTParams. (M)=236.8, Mem. (G)=82.02026.04 | 61.44 | — | — | |
| Ours‡Params. (M)=13.6, Mem. (G)=14.92026.04 | 61.44 | — | — | |
| Ours†Params. (M)=10.9, Mem. (G)=12.62026.04 | 61.41 | — | — | |
| OSCARModel Size=Base2020.04 | 61.23 | — | — | |
| OSCARModel Size=B2020.04 | 61.23 | — | — | |
| OSCARModel Size=Base2020.04 | 61.23 | — | — | |
| OscarVisual Encoder=BUTD-Res101, V&L Pretrain Data=6.5M, V&L Pretrain Epoch=1182021.07 | 61.23 | — | — | |
| LXMERT+CATTGPU Hours=960 (1080Ti), 624 (V100), Pre-training Data (Image / Text)=0.18M / 9.18M2021.03 | 61.17 | — | — | |
| Ours♠Params. (M)=13.1, Mem. (G)=15.2, Time (ms)=59.892026.04 | 60.93 | — | — | |
| Ours♡Params. (M)=10.4, Mem. (G)=12.2, Time (ms)=60.262026.04 | 60.85 | — | — | |
| MMN2020.04 | 60.83 | — | — | |
| MMN2021.01 | 60.83 | — | — | |
| MMN2021.04 | 60.83 | — | — | |
| SHERLParams. (M)=13.0, Mem. (G)=14.0, Time (ms)=92.492026.04 | 60.82 | — | — | |
| SHERLParams. (M)=13.0, Mem. (G)=14.02026.04 | 60.82 | — | — | |
| VL-T5Pretraining Images=180K2021.02 | 60.8 | — | — | |
| MMNVisual features=RCNN [2], Additional supervision=Program, Training images (M)=≈ 0.1, Training sentences (M)=≈ 152021.06 | 60.8 | 78.9 | 44.9 | |
| VL-T5Pre-training img data=VG, COCO (180k)2021.04 | 60.8 | — | — | |
| LSTArchitecture=CLIP-ViL2023.08 | 60.75 | — | — | |
| LSTParams. (M)=13.4, Mem. (G)=25.62026.04 | 60.75 | — | — | |
| LSTParams. (M)=13.4, Mem. (G)=25.62026.04 | 60.75 | — | — | |
| UniPTArchitecture=CLIP-ViL2023.08 | 60.72 | — | — | |
| UniPTParams. (M)=10.3, Mem. (G)=11.6, Time (ms)=106.382026.04 | 60.72 | — | — | |
| UniPTParams. (M)=10.3, Mem. (G)=11.62026.04 | 60.72 | — | — | |
| 12-in-12020.04 | 60.65 | — | — | |
| 12-in-12020.04 | 60.65 | — | — | |
| 12-in-12020.04 | 60.65 | — | — | |
| 12-in-12021.01 | 60.65 | — | — | |
| CLIP-ViLpVisual Encoder=CLIP-Res50, V&L Pretrain Data=9.2M, V&L Pretrain Epoch=202021.07 | 60.55 | — | — | |
| VL-BARTPretraining Images=180K2021.02 | 60.5 | — | — | |
| LXMERT2020.04 | 60.33 | — | — | |
| LXMERT2020.04 | 60.33 | — | — | |
| LXMERT2020.04 | 60.33 | — | — | |
| LXMERT2021.01 | 60.33 | — | — | |
| LXMERT2021.07 | 60.33 | — | — | |
| LXMERTPre-training img data=VG, COCO (180k)2021.04 | 60.33 | — | — | |
| LXMERTGPU Hours=960 (titan xp), Pre-training Data (Image / Text)=0.18M / 9.18M2021.03 | 60.3 | — | — | |
| LXMERTPretraining Images=180K2021.02 | 60.3 | — | — | |
| LXMERTVisual features=RCNN [2], Additional supervision=None, Training images (M)=~0.18, Training sentences (M)=~ 92021.06 | 60.3 | 77.8 | 45 | |
| LXMERTVisual Encoder=BUTD-Res101, V&L Pretrain Data=9.2M, V&L Pretrain Epoch=202021.07 | 60.3 | — | — | |
| LXMERT+GPU Hours=816 (1080Ti), Pre-training Data (Image / Text)=0.18M / 9.18M2021.03 | 59.94 | — | — | |
| LXMERT-S2021.07 | 59.12 | — | — | |
| Oracle transferVisual features=RCNN [2], Additional supervision=None, Training images (M)=~0.18, Training sentences (M)=≈ 12021.06 | 58.7 | 75.2 | 44.1 | |
| MCANVisual features=RCNN [2], Additional supervision=None, Training images (M)=≈ 0.1, Training sentences (M)=≈ 12021.06 | 58 | 75.9 | 42.2 | |
| BAN4Visual features=RCNN [2], Additional supervision=None, Training images (M)=≈ 0.1, Training sentences (M)=≈ 12021.06 | 57.1 | 76 | 40.4 | |
| MoVie2021.04 | 57.1 | — | — | |
| CLIP-BART + Full fine-tuningUpdated Params (%)=100.002021.12 | 52.5 | — | — | |
| CLIP-BART + Single AdapterUpdated Params (%)=4.182021.12 | 50.9 | — | — | |
| CLIP-BART + Single LORAUpdated Params (%)=5.932021.12 | 50 | — | — | |
| CLIP-BART + Single PromptUpdated Params (%)=2.002021.12 | 37.3 | — | — |