Image Captioning on MS-COCO
154.9CIDErOFA
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| OFAExample=18M, SCST finetuning=true2023.06 | 154.9 | — | — | 26.6 | — | — | — | — | — | |
| GIT2Example=12.9B, SCST finetuning=true2023.06 | 152.7 | — | — | 26.4 | — | — | — | — | — | |
| VALORExample=433.5M, SCST finetuning=true2023.06 | 152.5 | — | — | 25.7 | — | — | — | — | — | |
| GITExample=0.8B, SCST finetuning=true2023.06 | 151.1 | — | — | 26.3 | — | — | — | — | — | |
| COSAExample=415M, parameters=1.2B, SCST finetuning=true2023.06 | 150.6 | — | — | 27 | — | — | — | — | — | |
| BEiT-3Example=21M, SCST finetuning=false2023.06 | 147.6 | — | — | 25.4 | — | — | — | — | — | |
| BLIP-2Example=129M, SCST finetuning=false2023.06 | 145.8 | — | — | — | — | — | — | — | — | |
| CoCaExample=4.8B, SCST finetuning=false2023.06 | 143.6 | — | — | 24.7 | — | — | — | — | — | |
| SimVLMExample=1.8B, SCST finetuning=false2023.06 | 143.3 | — | — | 25.4 | — | — | — | — | — | |
| OSCARLModel scale=Large2020.04 | 140 | 41.7 | 30.6 | 24.5 | — | — | — | — | — | |
| FlamingoExample=2.3B, SCST finetuning=false2023.06 | 138.1 | — | — | — | — | — | — | — | — | |
| mPLUG-2Example=417M, SCST finetuning=false2023.06 | 137.7 | — | — | 23.7 | — | — | — | — | — | |
| OSCARBModel scale=Base2020.04 | 137.6 | 40.5 | 29.7 | 22.8 | — | — | — | — | — | |
| BLIPExample=129M, SCST finetuning=false2023.06 | 136.7 | — | — | — | — | — | — | — | — | |
| LEMONSupervision level=Fully Supervised2022.11 | 133.3 | 40.3 | 30.2 | — | — | — | — | — | — | |
| SoTASModel scale=Small2020.04 | 129.8 | 38.9 | 29.2 | 22.4 | — | — | — | — | — | |
| SoTABModel scale=Base2020.04 | 129.3 | 39.5 | 29.3 | 23.2 | — | — | — | — | — | |
| CapPa L/14Backbone Scale=L/14, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 125.8 | — | — | — | — | — | — | — | — | |
| CLIP L/14Backbone Scale=L/14, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 124.4 | — | — | — | — | — | — | — | — | |
| OscarSupervision level=Fully Supervised2022.11 | 123.7 | 36.5 | 30.3 | — | — | — | — | — | — | |
| CLIP* L/14Backbone Scale=L/14, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 123.2 | — | — | — | — | — | — | — | — | |
| CLIPBackbone Scale=B/16, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 118.4 | — | — | — | — | — | — | — | — | |
| CapPaBackbone Scale=B/16, Pre-training Batch Size=8k, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 117.9 | — | — | — | — | — | — | — | — | |
| CapBackbone Scale=B/16, Pre-training Batch Size=8k, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 117.5 | — | — | — | — | — | — | — | — | |
| UniVLPSupervision level=Fully Supervised2022.11 | 116.9 | 36.5 | 28.4 | — | — | — | — | — | — | |
| CLIP* (16k)Backbone Scale=B/16, Pre-training Batch Size=16k, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 116.3 | — | — | — | — | — | — | — | — | |
| CLIP* (8k)Backbone Scale=B/16, Pre-training Batch Size=8k, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 115.8 | — | — | — | — | — | — | — | — | |
| BUTDSupervision level=Fully Supervised2022.11 | 113.5 | 36.2 | 27 | — | 77.2 | — | 56.4 | — | — | |
| ClipCapSupervision level=Fully Supervised2022.11 | 113.1 | 33.5 | 27.5 | — | 74.7 | — | — | — | — | |
| VisInContextText=Rendered Image, ICL Tokens=2048, Shots=322024.06 | 101.3 | — | — | — | — | — | — | — | — | |
| Open-Flamingo MOEText=Raw Text, ICL Tokens=256, Shots=322024.06 | 98.2 | — | — | — | — | — | — | — | — | |
| VisInContextText=Rendered Image, ICL Tokens=2048, Shots=42024.06 | 94.2 | — | — | — | — | — | — | — | — | |
| CapDecSupervision level=Weakly or Unsupervised2022.11 | 91.8 | 26.4 | 25.1 | — | 69.2 | — | 51.8 | — | — | |
| Open-Flamingo MOEText=Raw Text, ICL Tokens=256, Shots=42024.06 | 90.5 | — | — | — | — | — | — | — | — | |
| VisInContextText=Rendered Image, ICL Tokens=2048, Shots=02024.06 | 84.4 | — | — | — | — | — | — | — | — | |
| Open-Flamingo MOEText=Raw Text, ICL Tokens=256, Shots=02024.06 | 82.3 | — | — | — | — | — | — | — | — | |
| FlamingoParams (#trainable)=9B, Text Backbone=AR, Image Backbone=-2024.12 | 79.4 | — | — | — | — | — | — | — | — | |
| FlamingoParams (B)=9, Text Gen Arch=AR, Image Gen Arch=-2025.05 | 79.4 | — | — | — | — | — | — | — | — | |
| FlamingoParams # trainable=9B, Text Backbone=AR, Image Backbone=-2026.05 | 79.4 | — | — | — | — | — | — | — | — | |
| VisualGPTTraining Data=1% images2021.02 | 75.8 | 24.3 | 21.9 | — | 67.1 | 48.6 | — | — | — | |
| OpenFlamingoParams (#trainable)=9B, Text Backbone=AR, Image Backbone=-2024.12 | 65.5 | — | — | — | — | — | — | — | — | |
| OpenFlamingoParams (B)=9, Text Gen Arch=AR, Image Gen Arch=-2025.05 | 65.5 | — | — | — | — | — | — | — | — | |
| OpenFlamingoParams # trainable=9B, Text Backbone=AR, Image Backbone=-2026.05 | 65.5 | — | — | — | — | — | — | — | — | |
| CM3LeonParams (#trainable)=7B, Text Backbone=AR, Image Backbone=AR2024.12 | 61.6 | — | — | — | — | — | — | — | — | |
| CM3LeonParams # trainable=7B, Text Backbone=AR, Image Backbone=AR2026.05 | 61.6 | — | — | — | — | — | — | — | — | |
| CM3LeonTrainable Params=7B2026.05 | 61.6 | — | — | — | — | — | — | — | — | |
| MudditParams (B)=1, Text Gen Arch=Discrete Diff., Image Gen Arch=Discrete Diff., Resolution=1024×10242025.05 | 60.1 | — | — | — | — | — | — | — | — | |
| MudditParams (B)=1, Text Gen Arch=Discrete Diff., Image Gen Arch=Discrete Diff., Resolution=512×5122025.05 | 59.9 | — | — | — | — | — | — | — | — | |
| D-DiTParams (#trainable)=2B, Text Backbone=Diffusion, Image Backbone=Diffusion, Input Resolution=512x5122024.12 | 56.2 | — | — | — | — | — | — | — | — | |
| D-DiTParams (B)=2, Text Gen Arch=Discrete Diff., Image Gen Arch=Diffusion, Resolution=512×5122025.05 | 56.2 | — | — | — | — | — | — | — | — | |
| DualDiff (512x512)Params # trainable=2B, Text Backbone=FM, Image Backbone=FM2026.05 | 56.2 | — | — | — | — | — | — | — | — | |
| DualDiffTrainable Params=2B, Training resolution=5122026.05 | 56.2 | — | — | — | — | — | — | — | — | |
| Kim et al. + unpairedTraining Data=1% images + 99% MS COCO unpaired images/texts, Supervision=Semi-supervised2021.02 | 55.2 | 18.7 | 20.7 | — | 63 | — | — | — | — | |
| FullFlowParams # trainable=130M, Text Backbone=FM, Image Backbone=FM2026.05 | 54.93 | — | — | — | — | — | — | — | — | |
| FullFlowTrainable Params=130M2026.05 | 54.93 | — | — | — | — | — | — | — | — | |
| Feng et al.Learning Setting=Unsupervised, Training Data=Tens of millions of unpaired images and captions2021.02 | 54.9 | 18.6 | 17.9 | — | 58.9 | — | — | — | — | |
| MAGICSupervision level=Weakly or Unsupervised2022.11 | 49.3 | 12.9 | 17.4 | — | 56.8 | — | 39.9 | — | — | |
| UniDiscParams (B)=1.4, Text Gen Arch=Discrete Diff., Image Gen Arch=Discrete Diff.2025.05 | 46.8 | — | — | — | — | — | — | — | — | |
| Kim et al.Training Data=1% images, Supervision=Semi-supervised2021.02 | 36 | 13.4 | 15.9 | — | 58.1 | — | — | — | — | |
| ZeroCapSupervision level=Weakly or Unsupervised2022.11 | 34.5 | 7 | 15.4 | — | 49.8 | — | 31.8 | — | — | |
| TransfusionParams (#trainable)=7B, Text Backbone=AR, Image Backbone=Diffusion2024.12 | 29 | — | — | — | — | — | — | — | — | |
| TransfusionParams (B)=7, Text Gen Arch=AR, Image Gen Arch=Diffusion2025.05 | 29 | — | — | — | — | — | — | — | — | |
| TransfusionParams # trainable=7B, Text Backbone=AR, Image Backbone=FM2026.05 | 29 | — | — | — | — | — | — | — | — | |
| TransfusionTrainable Params=7B2026.05 | 29 | — | — | — | — | — | — | — | — | |
| CapDecSource Training Dataset=Flickr30k2022.11 | 27.3 | 9.2 | 16.3 | — | 43.3 | — | 36.7 | — | — | |
| MAGICSource Training Dataset=Flickr30k2022.11 | 18.3 | 5.2 | 12.5 | — | 41.4 | — | 30.7 | — | — | |
| ChameleonParams (#trainable)=7B, Text Backbone=AR, Image Backbone=AR2024.12 | 18 | — | — | — | — | — | — | — | — | |
| ChameleonParams (B)=7, Text Gen Arch=AR, Image Gen Arch=AR2025.05 | 18 | — | — | — | — | — | — | — | — | |
| ChameleonParams # trainable=7B, Text Backbone=AR, Image Backbone=AR2026.05 | 18 | — | — | — | — | — | — | — | — | |
| ChameleonTrainable Params=7B2026.05 | 18 | — | — | — | — | — | — | — | — | |
| Gu et al.Learning Setting=Unsupervised, Training Data=Tens of millions of unpaired images and captions2021.02 | 17.7 | 5.4 | 13.2 | — | 46.2 | — | — | — | — | |
| RefineCap2021.09 | 1.272 | 37.8 | 28.3 | 0.225 | 80.2 | — | 0.58 | 64.5 | 49.9 | |
| Autoregressive Transformer2022.08 | 1.18 | 33.9 | — | — | — | — | 57 | — | — | |
| Bit DiffusionSampling steps=202022.08 | 1.15 | 34.7 | — | — | — | — | 58 | — | — | |
| Bit DiffusionSampling steps=402022.08 | 1.15 | 34.4 | — | — | — | — | 57 | — | — | |
| Bit DiffusionSampling steps=102022.08 | 1.13 | 34.5 | — | — | — | — | 57 | — | — | |
| Bridging2021.09 | 1.066 | 33 | 26.4 | — | — | — | 0.586 | — | — | |
| SCN-LSTM2021.09 | 1.041 | 33 | 25.7 | — | 72.8 | — | — | 56.6 | 43.3 | |
| Bit DiffusionSampling steps=52022.08 | 1 | 31.5 | — | — | — | — | 55 | — | — | |
| Skeleton Key2021.09 | 0.966 | 25.9 | 24.7 | 0.196 | 67.3 | — | 0.489 | 48.9 | 35.5 | |
| Att-CNN+LSTM2021.09 | 0.94 | 31 | 26 | — | 74 | — | — | 56 | 42 | |
| BLIP-B#param=252M, #data=129M, protocol=WT2022.06 | — | 39.7 | — | — | — | — | — | — | — | |
| BLIP-L#param=473M, #data=129M, protocol=WT2022.06 | — | 40.4 | — | — | — | — | — | — | — | |
| CLIP-VIL#param=>459M, #data=400M, protocol=WT2022.06 | — | 40.2 | — | — | — | — | — | — | — | |
| CoCa#param=472M, #data=60.6M, protocol=WT2022.06 | — | 43.5 | — | — | — | — | — | — | — | |
| LSTM-C2021.09 | — | — | — | — | — | — | 0.23 | — | — | |
| OSCAR-B#param=154M, #data=6.5M, region features=true, protocol=WT2022.06 | — | 36.5 | — | — | — | — | — | — | — | |
| OSCAR-L#param=384M, #data=6.5M, region features=true, protocol=WT2022.06 | — | 37.4 | — | — | — | — | — | — | — | |
| OSCAR-L#param=384M, #data=6.5M, region features=true, CIDEr optimization=true, protocol=WT2022.06 | — | 41.7 | — | — | — | — | — | — | — | |
| SemAttn2021.09 | — | 30.4 | 24.3 | — | 70.9 | — | — | 53.7 | 40.2 | |
| SimVLM#param=632M, #data=1.8B, protocol=WT2022.06 | — | 40.6 | — | — | — | — | — | — | — | |
| Uni-Perceiver-B#param=124M, #data=44.1M, protocol=Zero-shot (WT)2022.06 | — | 32 | — | — | — | — | — | — | — | |
| Uni-Perceiver-B#param=124M, #data=44.1M, protocol=Prompt Tuning (PT1%)2022.06 | — | 36.8 | — | — | — | — | — | — | — | |
| Uni-Perceiver-B#param=124M, #data=44.1M, protocol=Fine-tuning (FT100%)2022.06 | — | 37.3 | — | — | — | — | — | — | — | |
| Uni-Perceiver-B + Conditional MoEs#param=167M, #data=44.1M, protocol=Zero-shot (WT)2022.06 | — | 33.2 | — | — | — | — | — | — | — | |
| Uni-Perceiver-B + Conditional MoEs#param=167M, #data=44.1M, protocol=Prompt Tuning (PT1%)2022.06 | — | 38.6 | — | — | — | — | — | — | — | |
| Uni-Perceiver-B + Conditional MoEs#param=167M, #data=44.1M, protocol=Fine-tuning (FT100%)2022.06 | — | 39.2 | — | — | — | — | — | — | — | |
| Uni-Perceiver-L#param=354M, #data=44.1M, protocol=Zero-shot (WT)2022.06 | — | 35.3 | — | — | — | — | — | — | — | |
| Uni-Perceiver-L#param=354M, #data=44.1M, protocol=Prompt Tuning (PT1%)2022.06 | — | 39.3 | — | — | — | — | — | — | — | |
| Uni-Perceiver-L#param=354M, #data=44.1M, protocol=Fine-tuning (FT100%)2022.06 | — | 40.5 | — | — | — | — | — | — | — |