Visual Question Answering on DAQUAR-ALL full (test)
50.2AccuracyHuman
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Human2015.11 | 50.2 | 50.8 | 67.3 | |
| Human Baseline2016.03 | 50.2 | 50.82 | 62.27 | |
| SAN(2, LSTM)Layers=2, Text Encoder=LSTM2015.11 | 29.3 | 34.9 | 68.1 | |
| SAN(2, CNN)Layers=2, Text Encoder=CNN2015.11 | 29.3 | 35.1 | 68.6 | |
| Yang et al.2016.03 | 29.3 | 35.1 | 68.6 | |
| A+C+Selected-K-LSTMknowledge_selection=true2016.03 | 29.23 | 35.37 | 68.72 | |
| SAN(1, CNN)Layers=1, Text Encoder=CNN2015.11 | 29.2 | 35.1 | 67.8 | |
| Att+Cap+Know-LSTMincludes_attention=true, includes_captions=true, includes_knowledge=true2016.03 | 29.16 | 35.3 | 68.66 | |
| Noh et al.2016.03 | 28.98 | 34.8 | 67.81 | |
| SAN(1, LSTM)Layers=1, Text Encoder=LSTM2015.11 | 28.9 | 34.7 | 68.5 | |
| Att+Cap-LSTMincludes_captions=true2016.03 | 27.04 | 33.4 | 67.65 | |
| Att+Know-LSTMincludes_knowledge=true2016.03 | 24.89 | 31.27 | 66.11 | |
| Att-LSTM2016.03 | 24.27 | 30.41 | 62.29 | |
| Cap+Know-LSTMincludes_captions=true, includes_knowledge=true2016.03 | 23.91 | 30.64 | 65.01 | |
| VggNet+ft-LSTMbackbone=VggNet, fine-tuned=true2016.03 | 23.75 | 30.22 | 63.66 | |
| IMG-CNNSource=CNN2015.11 | 23.4 | 29.6 | 63 | |
| Ma et al.2016.03 | 23.4 | 29.59 | 62.95 | |
| VggNet-LSTMbackbone=VggNet2016.03 | 23.13 | 30.01 | 63.61 | |
| Language + IMGSource=Ask-Your-Neurons2015.11 | 21.7 | 28 | 65 | |
| Askneuron2016.03 | 19.43 | 25.28 | 62 | |
| LanguageSource=Ask-Your-Neurons2015.11 | 19.1 | 25.2 | 65.1 | |
| Multi-World2015.11 | 7.9 | 11.9 | 38.8 |