Visual Question Answering (Multiple-choice) on VQA 1.0 (test-dev)
70.04Accuracy (All)Dual-MFA
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| Dual-MFA2017.11 | 70.04 | — | — | — | |
| RelAttattention_type=Semantic attention2018.05 | 69.6 | 83.58 | 64.65 | 38.56 | |
| Dual-MLBtype=baseline2017.11 | 69.55 | — | — | — | |
| Fukui et al. 2016 ResNet*Model Architecture=MCB, Backbone=ResNet2016.06 | 69.1 | — | — | — | |
| DANBackbone=ResNet2016.11 | 69.1 | — | — | — | |
| MCB2017.11 | 69.1 | — | — | — | |
| MCBattention_type=Visual attention2018.05 | 68.6 | — | — | — | |
| MCBBackbone=ResNet2016.11 | 68.6 | — | — | — | |
| Ours ResNetBackbone=ResNet, Training Strategy=Multi-step training + Early-stopping2016.06 | 67.7 | 81.9 | 61.5 | 41.1 | |
| RAUBackbone=ResNet2016.11 | 67.7 | — | — | — | |
| VQA-Machine2017.11 | 67.7 | — | — | — | |
| DANBackbone=VGG2016.11 | 67 | — | — | — | |
| MRNBackbone=Residual, Target Answers=3k2016.06 | 66.33 | 82.41 | 58.4 | 39.57 | |
| MRNBackbone=ResNet2016.11 | 66.2 | — | — | — | |
| MRNattention_type=Visual attention2018.05 | 66.15 | 82.3 | 58.16 | 40.45 | |
| MRNBackbone=Residual, Target Answers=2k2016.06 | 66.15 | 82.3 | 58.16 | 40.45 | |
| MRN2017.11 | 66.15 | — | — | — | |
| Ours_FULLBackbone=VGG-16, Training Strategy=Multi-step training + Early-stopping2016.06 | 66.1 | 81.5 | 58.9 | 39.5 | |
| Lu et al. 2016 ResNet*Model Architecture=HieCoAtt, Backbone=ResNet2016.06 | 65.8 | 79.7 | 59.8 | 40 | |
| Ours_MSBackbone=VGG-16, Training Strategy=Multi-step training2016.06 | 65.8 | 81.3 | 58.7 | 38.9 | |
| HieCoAttBackbone=ResNet2016.11 | 65.8 | — | — | — | |
| HieCoAtt2017.11 | 65.8 | — | — | — | |
| MRNBackbone=Residual, Target Answers=1k2016.06 | 65.62 | 82.39 | 57.15 | 39.65 | |
| QRUattention_type=Visual attention2018.05 | 65.43 | 82.24 | 57.12 | 38.69 | |
| QRU2017.11 | 65.43 | — | — | — | |
| MRNBackbone=VGG, Target Answers=2k2016.06 | 65.27 | 82.12 | 56.39 | 40.84 | |
| MRNBackbone=VGG, Target Answers=3k2016.06 | 65.09 | 82.42 | 55.93 | 40.13 | |
| MLANattention_type=Semantic attention2018.05 | 64.8 | — | — | — | |
| MRNBackbone=VGG, Target Answers=1k2016.06 | 64.79 | 82.55 | 55.23 | 39.93 | |
| FDA2016.04 | 64.01 | 81.5 | 54.72 | 39 | |
| FDAattention_type=No attention2018.05 | 64.01 | 81.5 | 54.72 | 39 | |
| FDA2016.06 | 64.01 | 81.5 | 54.72 | 39 | |
| FDA2017.11 | 64.01 | — | — | — | |
| Ours_SSBackbone=VGG-16, Training Strategy=Single-step training2016.06 | 64 | 78.4 | 57.5 | 38.2 | |
| Deep Q+I2016.06 | 62.86 | 80.88 | 53.14 | 37.78 | |
| Antol et al. 2015Model Architecture=VQA Baseline2016.06 | 62.7 | 80.5 | 53 | 38.2 | |
| VQA team2016.11 | 62.7 | — | — | — | |
| Andreas et al. 2016bModel Architecture=NMN2016.06 | 62.5 | 80.8 | 52.2 | 38.9 | |
| DPPnet2016.11 | 62.5 | — | — | — | |
| DPPnet2016.04 | 62.48 | 80.79 | 52.16 | 38.94 | |
| DPPnetattention_type=No attention2018.05 | 62.48 | 80.79 | 52.16 | 38.94 | |
| DPPnet2016.06 | 62.48 | 80.79 | 52.16 | 38.94 | |
| DPPnet2017.11 | 62.48 | — | — | — | |
| Region Sel.2017.11 | 62.44 | — | — | — | |
| Shih, Singh, and Hoiem 20162016.06 | 62.4 | 77.6 | 55.8 | 34.3 | |
| Noh, Seo, and Han 2016Model Architecture=DPPnet2016.06 | 61.7 | 76.7 | 54.4 | 37.1 | |
| iBOWIMG2016.11 | 61.7 | — | — | — | |
| iBOWIMG2016.04 | 61.68 | 76.68 | 54.44 | 38.94 | |
| iBOWING2017.11 | 61.68 | — | — | — | |
| WR2016.04 | 60.96 | — | — | — | |
| Bow Q+I (Antol et al. 2015)Input=Question + Image, Feature=Bag-of-Words2016.06 | 59 | 75.6 | 50.3 | 34.4 | |
| Q+I2016.04 | 58.97 | 75.59 | 50.33 | 34.35 | |
| Q+I2016.06 | 58.97 | 75.59 | 50.33 | 34.35 | |
| LSTM Q+I (Antol et al. 2015)Input=Question + Image, Feature=LSTM2016.06 | 57.2 | 79 | 43.4 | 35.8 | |
| LSTM Q+I2016.04 | 57.17 | 78.95 | 43.41 | 35.8 | |
| LSTM Q+Iattention_type=No attention2018.05 | 57.17 | 78.85 | 43.41 | 35.8 | |
| LSTM Q+Iquestion_model=LSTM2016.06 | 57.17 | 78.95 | 43.41 | 35.8 | |
| LSTM Q+I2017.11 | 57.17 | — | — | — | |
| LSTM Q (Antol et al. 2015)Input=Question only, Feature=LSTM2016.06 | 54.8 | 78.2 | 38.8 | 36.8 | |
| LSTM Qquestion_model=LSTM2016.06 | 54.75 | 78.22 | 38.78 | 36.82 | |
| BOW Q (Antol et al. 2015)Input=Question only, Feature=Bag-of-Words2016.06 | 53.7 | 75.7 | 38.6 | 37.1 | |
| Question2016.04 | 53.68 | 75.71 | 38.64 | 37.05 | |
| Question2016.06 | 53.68 | 75.71 | 38.64 | 37.05 | |
| Image2016.04 | 30.53 | 69.87 | 3.76 | 0.45 | |
| Image2016.06 | 30.53 | 69.87 | 3.76 | 0.45 | |
| I (Antol et al. 2015)Input=Image only2016.06 | 30.5 | 69.9 | 3.8 | 0.5 |