Open-Ended Visual Question Answering on VQA 1.0 (test-dev)
66.7Overall AccuracyEnsemble of 7 Att. models
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Ensemble of 7 Att. modelsTraining Split=train+val, Ensemble=7 models2016.06 | 66.7 | 83.4 | 39.8 | 58.5 | |
| Dual-MFA2017.11 | 66.01 | 83.59 | 40.18 | 56.84 | |
| RelAttattention_type=Semantic attention2018.05 | 65.69 | 83.55 | 36.92 | 56.94 | |
| MCB + Att. + GloVe + GenomeTraining Split=train+val, Use Attention=Yes, Use GloVe=Yes, Use Visual Genome=Yes2016.06 | 65.4 | 82.3 | 37.2 | 57.4 | |
| Dual-MLBtype=baseline2017.11 | 65.12 | 83.32 | 39.96 | 55.26 | |
| MCB + Att. + GenomeTraining Split=train+val, Use Attention=Yes, Use Visual Genome=Yes2016.06 | 65.1 | 81.7 | 38.2 | 57 | |
| Naver Labs (challenge 2nd)Training Split=train+val2016.06 | 64.9 | 83.5 | 39.8 | 54.8 | |
| MLB2017.11 | 64.89 | 84.13 | 37.85 | 54.57 | |
| MCB + Att. + GloVeTraining Split=train+val, Use Attention=Yes, Use GloVe=Yes2016.06 | 64.7 | 82.5 | 37.6 | 55.6 | |
| Fukui et al. 2016 ResNet*Model Architecture=MCB, Backbone=ResNet2016.06 | 64.7 | 82.5 | 37.6 | 55.6 | |
| MCB2017.11 | 64.7 | 82.5 | 37.6 | 55.6 | |
| MLANattention_type=Semantic attention2018.05 | 64.6 | 83.8 | 40.2 | 53.7 | |
| MLBattention_type=Visual attention2018.05 | 64.53 | 83.41 | 37.82 | 54.43 | |
| DANBackbone=ResNet2016.11 | 64.3 | 83 | 39.1 | 53.9 | |
| MCBattention_type=Visual attention2018.05 | 64.2 | 82.2 | 37.7 | 54.8 | |
| MCB + Att.Training Split=train+val, Use Attention=Yes, Use Visual Genome=No2016.06 | 64.2 | 82.2 | 37.7 | 54.8 | |
| MCBBackbone=ResNet2016.11 | 64.2 | 82.2 | 37.7 | 54.8 | |
| Ours ResNetBackbone=ResNet, Training Strategy=Multi-step training + Early-stopping2016.06 | 63.3 | 81.9 | 39 | 53 | |
| RAUBackbone=ResNet2016.11 | 63.3 | 81.9 | 39 | 53 | |
| VQA-Machine2017.11 | 63.1 | 81.5 | 38.4 | 53 | |
| MCB + GenomeTraining Split=train+val, Use Attention=No, Use Visual Genome=Yes2016.06 | 62.3 | 81.7 | 36.6 | 51.5 | |
| DANBackbone=VGG2016.11 | 62 | 82.1 | 38.2 | 50.2 | |
| HieCoAttTraining Split=train+val2016.06 | 61.8 | 79.7 | 38.7 | 51.7 | |
| Lu et al. 2016 ResNet*Model Architecture=HieCoAtt, Backbone=ResNet2016.06 | 61.8 | 79.7 | 38.7 | 51.7 | |
| HieCoAttBackbone=ResNet2016.11 | 61.8 | 79.7 | 38.7 | 51.7 | |
| HieCoAtt2017.11 | 61.8 | 79.7 | 38.7 | 51.7 | |
| Kim et al. 2016 ResNet*Model Architecture=MLB, Backbone=ResNet2016.06 | 61.7 | 82.3 | 38.8 | 49.3 | |
| MRNBackbone=ResNet2016.11 | 61.7 | 82.3 | 38.8 | 49.3 | |
| MRNBackbone=Residual, Target Answers=2k2016.06 | 61.68 | 82.28 | 38.82 | 49.25 | |
| MRNattention_type=Visual attention2018.05 | 61.68 | 82.28 | 38.82 | 49.25 | |
| MRN2017.11 | 61.68 | 82.28 | 38.82 | 49.25 | |
| MRNBackbone=Residual, Target Answers=3k2016.06 | 61.47 | 82.28 | 39.09 | 48.76 | |
| MRNBackbone=Residual, Target Answers=1k2016.06 | 61.45 | 82.36 | 38.4 | 48.81 | |
| Ours_FULLBackbone=VGG-16, Training Strategy=Multi-step training + Early-stopping2016.06 | 61.3 | 81.5 | 37 | 49.6 | |
| Ours_MSBackbone=VGG-16, Training Strategy=Multi-step training2016.06 | 61 | 81.3 | 36.8 | 49.1 | |
| MCBTraining Split=train+val, Use Attention=No, Use Visual Genome=No2016.06 | 60.8 | 81.2 | 35.1 | 49.3 | |
| MRNBackbone=VGG, Target Answers=2k2016.06 | 60.77 | 82.1 | 39.11 | 47.46 | |
| QRUattention_type=Visual attention2018.05 | 60.72 | 82.29 | 37.02 | 47.67 | |
| QRU2017.11 | 60.72 | 82.29 | 37.02 | 47.67 | |
| MRNBackbone=VGG, Target Answers=3k2016.06 | 60.68 | 82.4 | 38.69 | 47.1 | |
| MRNBackbone=VGG, Target Answers=1k2016.06 | 60.53 | 82.53 | 38.34 | 46.78 | |
| DMN+2016.06 | 60.3 | 80.5 | 36.8 | 48.3 | |
| DMN+attention_type=No attention2018.05 | 60.3 | 80.5 | 36.8 | 48.3 | |
| DMN+Training Split=train+val2016.06 | 60.3 | 80.5 | 36.8 | 48.3 | |
| Xiong, Merity, and Socher 2016Model Architecture=DMN+2016.06 | 60.3 | 80.5 | 36.8 | 48.3 | |
| DMN+2016.11 | 60.3 | 80.5 | 36.8 | 48.3 | |
| DMN+2017.11 | 60.3 | 80.5 | 36.8 | 48.3 | |
| D-NMNTraining Split=train+val2016.06 | 59.4 | 81.1 | 38.6 | 45.5 | |
| FDA2016.06 | 59.24 | 81.14 | 36.16 | 45.77 | |
| FDAattention_type=No attention2018.05 | 59.24 | 81.14 | 36.16 | 45.77 | |
| FDA2017.11 | 59.24 | 81.14 | 36.16 | 45.77 | |
| FDATraining Split=train+val2016.06 | 59.2 | 81.1 | 36.2 | 45.8 | |
| AMATraining Split=train+val2016.06 | 59.2 | 81 | 38.4 | 45.2 | |
| Wu et al. 20162016.06 | 59.2 | 81 | 38.4 | 45.2 | |
| ACK2016.11 | 59.2 | 81 | 38.4 | 45.2 | |
| ACK2016.06 | 59.17 | 81.01 | 38.42 | 45.23 | |
| AMAattention_type=Semantic attention2018.05 | 59.17 | 81.01 | 38.42 | 45.23 | |
| Ours_SSBackbone=VGG-16, Training Strategy=Single-step training2016.06 | 59 | 78.3 | 37.1 | 47.6 | |
| SAN2016.06 | 58.7 | 79.3 | 36.6 | 46.1 | |
| SANattention_type=Visual attention2018.05 | 58.7 | 79.3 | 36.6 | 46.1 | |
| SANTraining Split=train+val2016.06 | 58.7 | 79.3 | 36.6 | 46.1 | |
| Yang et al. 2016Model Architecture=SAN2016.06 | 58.7 | 79.3 | 36.6 | 46.1 | |
| SAN2016.11 | 58.7 | 79.3 | 36.6 | 46.1 | |
| SAN2017.11 | 58.7 | 79.3 | 36.6 | 46.1 | |
| NMNTraining Split=train+val2016.06 | 58.6 | 81.2 | 38 | 44 | |
| NMN2016.11 | 58.6 | 81.2 | 38 | 44 | |
| AYNTraining Split=train+val2016.06 | 58.4 | 78.4 | 36.4 | 46.3 | |
| Deep Q+I2016.06 | 58.02 | 80.87 | 36.46 | 43.4 | |
| SMemTraining Split=train+val2016.06 | 58 | 80.9 | 37.3 | 43.1 | |
| SMemattention_type=Visual attention2018.05 | 57.99 | 80.87 | 37.32 | 43.12 | |
| SMem2017.11 | 57.99 | 80.87 | 37.32 | 43.12 | |
| D-NMN2016.06 | 57.9 | 80.5 | 37.4 | 43.1 | |
| Andreas et al. 2016bModel Architecture=NMN2016.06 | 57.9 | 80.5 | 37.4 | 43.1 | |
| VQA teamTraining Split=train+val2016.06 | 57.8 | 80.5 | 36.8 | 43.1 | |
| Antol et al. 2015Model Architecture=VQA Baseline2016.06 | 57.8 | 80.5 | 36.8 | 43.1 | |
| VQA team2016.11 | 57.8 | 80.5 | 36.8 | 43.1 | |
| V2Lattention_type=Semantic attention2018.05 | 57.46 | 78.9 | 36.11 | 40.07 | |
| Andreas et al. 2016aModel Architecture=NMN2016.06 | 57.3 | 79.7 | 37.3 | 39.3 | |
| DPPnet2016.06 | 57.22 | 80.71 | 37.24 | 41.69 | |
| DPPnetattention_type=No attention2018.05 | 57.22 | 80.71 | 37.24 | 41.69 | |
| DPPnet2017.11 | 57.22 | 80.71 | 37.24 | 41.69 | |
| DPPnetTraining Split=train+val2016.06 | 57.2 | 80.7 | 37.2 | 41.7 | |
| Noh, Seo, and Han 2016Model Architecture=DPPnet2016.06 | 57.2 | 80.7 | 37.2 | 41.7 | |
| DPPnet2016.11 | 57.2 | 80.7 | 37.2 | 41.7 | |
| iBOWING2017.11 | 55.72 | 76.55 | 35.03 | 42.62 | |
| iBOWIMGTraining Split=train+val2016.06 | 55.7 | 76.5 | 35 | 42.6 | |
| Zhou et al. 2015Model Architecture=iBOWIMG2016.06 | 55.7 | 76.6 | 35 | 42.6 | |
| iBOWIMG2016.11 | 55.7 | 76.5 | 35 | 42.6 | |
| LSTM Q+Iquestion_model=LSTM2016.06 | 53.74 | 78.94 | 35.24 | 36.42 | |
| LSTM Q+Iattention_type=No attention2018.05 | 53.74 | 78.94 | 35.24 | 36.42 | |
| LSTM Q+I2017.11 | 53.74 | 78.94 | 35.24 | 36.42 | |
| LSTM Q+I (Antol et al. 2015)Input=Question + Image, Feature=LSTM2016.06 | 53.7 | 78.9 | 35.2 | 36.4 | |
| Q+I2016.06 | 52.64 | 75.55 | 33.67 | 37.37 | |
| Bow Q+I (Antol et al. 2015)Input=Question + Image, Feature=Bag-of-Words2016.06 | 52.6 | 75.6 | 33.7 | 37.4 | |
| LSTM Q (Antol et al. 2015)Input=Question only, Feature=LSTM2016.06 | 48.8 | 78.2 | 35.7 | 26.6 | |
| LSTM Qquestion_model=LSTM2016.06 | 48.76 | 78.2 | 35.68 | 26.59 | |
| BOW Q (Antol et al. 2015)Input=Question only, Feature=Bag-of-Words2016.06 | 48.1 | 75.7 | 36.7 | 27.1 | |
| Question2016.06 | 48.09 | 75.66 | 36.7 | 27.14 | |
| Image2016.06 | 28.13 | 64.01 | 0.42 | 3.77 | |
| I (Antol et al. 2015)Input=Image only2016.06 | 28.1 | 64 | 0.4 | 3.8 |