Visual Question Answering on abstract scenes (test)
74.37Overall Accuracy (MC)Graph VQA
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Graph VQAmodel_version=full model2016.09 | 74.37 | 79.74 | 68.31 | 74.97 | 70.42 | 81.26 | 56.28 | 76.47 | |
| U. Tokyo MILensemble=true2016.09 | 71.18 | 79.59 | 67.93 | 56.19 | 69.73 | 80.7 | 62.08 | 58.82 | |
| LSTM with global image featuresfeatures=global image features2016.09 | 69.21 | 77.46 | 66.65 | 52.9 | 65.02 | 77.45 | 56.41 | 52.54 | |
| Multimodal residual learning2016.09 | 67.99 | 79.08 | 61.99 | 52.57 | 62.56 | 79.1 | 48.9 | 51.6 | |
| LSTM blindtype=blind (language only)2016.09 | 61.41 | 76.9 | 49.19 | 49.65 | 57.19 | 76.88 | 38.79 | 49.55 | |
| Zhang et al.evaluation scope=yes/no only2016.09 | 35.25 | 79.14 | — | — | 35.25 | 79.14 | — | — |