Fact-based Visual Question Answering on FVQA 1.0 (test)
87.3WUPS@0.0 (Top-1)Human
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| HumanType=Human evaluation2016.06 | 87.3 | — | — | |
| KB-query model (gt QQmapping)Question-Query mapping=ground truth2016.06 | 73.98 | 78.67 | 79.98 | |
| KB-query model (top-3-QQmaping)Question-Query mapping=top-3 predicted2016.06 | 72.34 | 77.52 | 78.69 | |
| Hie-Question+Image+Pre-VQAInput Modalities=Question + Image, Architecture=Hierarchical, Use Pre-VQA=true2016.06 | 71.51 | 82.71 | 89.01 | |
| Hie-Question+ImageInput Modalities=Question + Image, Architecture=Hierarchical2016.06 | 66.11 | 78.9 | 86.32 | |
| KB-query model (top-1-QQmaping)Question-Query mapping=top-1 predicted2016.06 | 64.96 | 69.57 | 70.64 | |
| LSTM-Question+Image+Pre-VQAInput Modalities=Question + Image, Architecture=LSTM, Use Pre-VQA=true2016.06 | 63.42 | 76.63 | 84.94 | |
| LSTM-Question+ImageInput Modalities=Question + Image, Architecture=LSTM2016.06 | 61.86 | 74.45 | 83.54 | |
| SVM-Question+ImageInput Modalities=Question + Image, Classifier=SVM2016.06 | 59.97 | 73.39 | 81.93 | |
| SVM-ImageInput Modality=Image, Classifier=SVM2016.06 | 59.64 | 73.3 | 81.77 | |
| LSTM-ImageInput Modality=Image, Architecture=LSTM2016.06 | 59.53 | 73.83 | 83.53 | |
| SVM-QuestionInput Modality=Question, Classifier=SVM2016.06 | 56.88 | 68.45 | 76.93 | |
| LSTM-QuestionInput Modality=Question, Architecture=LSTM2016.06 | 51.45 | 65.65 | 75.76 |