Image Captioning on Flickr 8k
76BLEU-1Att-GT+LSTM
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Att-GT+LSTMAttribute Labels=Ground Truth2016.03 | 76 | 57 | 41 | 29 | 12.52 | |
| Att-RegionCNN+LSTMAttribute Labels=Predicted (Region CNN)2016.03 | 74 | 54 | 38 | 27 | 12.6 | |
| Att-SVM+LSTMAttribute Labels=Predicted (SVM)2016.03 | 73 | 53 | 38 | 26 | 12.63 | |
| Att-GlobalCNN+LSTMAttribute Labels=Predicted (Global CNN)2016.03 | 72 | 53 | 38 | 27 | 12.63 | |
| Human2014.11 | 70 | — | — | — | — | |
| Xu et al. (Hard-Attention)Method Name=Hard-Attention2016.03 | 67 | 46 | 31 | 21 | — | |
| Google(NIC)Method Name=NIC2016.03 | 66 | 42 | 27 | 18 | — | |
| VggNet+ft+LSTMBackbone=VggNet, Feature Processing=Fine-tuning2016.03 | 64 | 43 | 30 | 20 | 14.69 | |
| NIC2014.11 | 63 | — | — | — | — | |
| Karpathy & Li (NeuralTalk)Method Name=NeuralTalk2016.03 | 63 | 45 | 32 | 23 | — | |
| m-RNN2014.11 | 58 | — | — | — | — | |
| SOTA2014.11 | 58 | — | — | — | — | |
| VggNet+LSTMBackbone=VggNet, Feature Processing=Standard2016.03 | 58 | 38 | 25 | 16 | 15.71 | |
| VggNet-PCA+LSTMBackbone=VggNet, Feature Processing=PCA2016.03 | 58 | 38 | 25 | 16 | 16.07 | |
| Mao et al. (m-Rnn-AlexNet)Backbone=AlexNet, Method Name=m-Rnn2016.03 | 57 | 39 | 26 | 17 | 24.39 | |
| MNLMbackbone=OxfordNet2014.11 | 51 | — | — | — | — | |
| Tri5Sem2014.11 | 48 | — | — | — | — | |
| Chen & Zintick (Mind's Eye)Method Name=Mind's Eye2016.03 | — | — | — | 14 | 15.1 |