Image Captioning on COCO Captions (test)
123.7CIDErOSCAR
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| OSCARParams=110M, Architecture Type=Single-stream2021.06 | 123.7 | 36.5 | — | — | |
| AoANetOptimization=Cross-Entropy, Model Type=Single2019.09 | 119.8 | 37.2 | 28.4 | 21.3 | |
| VLPOptimization=Cross-Entropy, Model Type=Single, Pre-training strategy=seq2seq only2019.09 | 117.7 | 36.5 | 28.4 | 21.3 | |
| E2E-VLPParams=94M, Architecture Type=Our Model2021.06 | 117.3 | 36.2 | — | — | |
| Unified VLPOptimization=Cross-Entropy, Model Type=Single, Pre-training strategy=joint2019.09 | 116.9 | 36.5 | 28.4 | 21.2 | |
| VLPParams=110M, Architecture Type=Single-stream2021.06 | 116.9 | 36.5 | — | — | |
| VLPOptimization=Cross-Entropy, Model Type=Single, Pre-training strategy=bidirectional only2019.09 | 116.5 | 36.1 | 28.3 | 21.2 | |
| GCN-LSTMOptimization=Cross-Entropy, Model Type=Single, Feature Type=sem2019.09 | 116.3 | 36.8 | 27.9 | 20.9 | |
| GCN-LSTMOptimization=Cross-Entropy, Model Type=Single, Feature Type=spa2019.09 | 115.6 | 36.5 | 27.8 | 20.8 | |
| VLPOptimization=Cross-Entropy, Model Type=Single, Pre-training strategy=none (baseline)2019.09 | 114.3 | 35.5 | 28.2 | 21 | |
| BUTDOptimization=Cross-Entropy, Model Type=Single2019.09 | 113.5 | 36.2 | 27 | 20.3 | |
| Openai CLIPBackbone=ViT-L/142024.11 | 113.5 | — | — | — | |
| SuperClassBackbone=ViT-L/162024.11 | 113 | — | — | — | |
| CLIP,reimpl.Backbone=ViT-L/162024.11 | 112.6 | — | — | — | |
| NBTOptimization=Cross-Entropy, Model Type=Single, Input features=with BBox2019.09 | 107.2 | 34.7 | 27.1 | 20.1 |