Image Captioning on COCO (test) with Extended Metrics
140.9CIDErVinVL
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| VinVLtraining_mode=supervised2021.11 | 140.9 | 0.41 | 0.311 | 0.252 | 0.83 | 1,125 | 0.779 | 0.78 | |
| OscarData=4.1M2023.01 | 137.6 | 40.5 | 29.7 | 22.8 | — | — | — | — | |
| VinVL*#Param=112M, Data=3.17M2023.01 | 136.5 | 39.6 | 30.4 | 24.4 | — | — | — | — | |
| GIVL#Param=112M, Data=3.17M2023.01 | 135.1 | 39.6 | 30.3 | 24.3 | — | — | — | — | |
| SimVLM-base#Param=1.8B2023.01 | 134.8 | 39 | 32.9 | 24 | — | — | — | — | |
| CLIP-VLtraining_mode=supervised, architecture=transformer2021.11 | 134.2 | 0.402 | 0.297 | 0.238 | 0.82 | 2,464 | 0.851 | 0.77 | |
| CLIP-ViL#Param=178M2023.01 | 134.2 | 40.2 | 29.7 | 23.8 | — | — | — | — | |
| VLP2023.01 | 129.8 | 39.5 | 29.3 | 22.4 | — | — | — | — | |
| Unimo-Large#Param=300M2023.01 | 127.7 | 39.6 | — | — | — | — | — | — | |
| BUTD2023.01 | 120.1 | 36.3 | 27.7 | 21.4 | — | — | — | — | |
| VL-T5#Param=224M2023.01 | 116.5 | — | — | — | — | — | — | — | |
| ClipCaptraining_mode=supervised, decoder=fine-tuned GPT-22021.11 | 108.35 | 0.3215 | 0.271 | 0.2012 | 0.81 | 1,650 | 0.664 | 0.77 | |
| ZeroCaptraining_mode=zero-shot2021.11 | 14.6 | 0.026 | 0.115 | 0.055 | 0.79 | 8,681 | 1 | 0.87 |