Image Captioning on MSCOCO (BLEU, METEOR, CIDEr)
31.4BLEU@4TipCap
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| TipCapEncoder=ViT-L/14, Decoder=GPT2-large2025.12 | 31.4 | 73.3 | 26.9 | 106.6 | |
| IFCapEncoder=ViT-B/32, Decoder=GPT2-base2025.12 | 30.8 | — | 20.3 | 108 | |
| ArcSinEncoder=ViT-L/14, Decoder=T5-base2025.12 | 30.3 | — | — | 99.6 | |
| TOMCap (large)Encoder=ViT-L/16, Decoder=GPT2-large2025.12 | 30.2 | 73.9 | 26.6 | 108.3 | |
| ICSDEncoder=ViT-B/32, Decoder=BERT-base2025.12 | 29.9 | — | 25.4 | 96.6 | |
| SynTICEncoder=ViT-B/32, Decoder=Transformer2025.12 | 29.9 | — | 25.8 | 101.1 | |
| CLOSEEncoder=ViT-L/14, Decoder=T5-base2025.12 | 29.5 | — | 25.7 | 97.8 | |
| TOMCapEncoder=ViT-L/16, Decoder=GPT2-base2025.12 | 28.4 | 72.7 | 25.8 | 103.4 | |
| TOMCap (retrieval only)Encoder=ViT-L/16, Decoder=GPT2-base2025.12 | 28.2 | 72.4 | 25.2 | 101.6 | |
| EntroCapEncoder=ViT-B/32, Decoder=GPT2-base2025.12 | 27.6 | — | 25.3 | 94.3 | |
| ViECapEncoder=ViT-B/32, Decoder=GPT2-base2025.12 | 27.2 | — | 24.8 | 92.9 | |
| MeaCapInvLMEncoder=ViT-B/32, Decoder=GPT2-base2025.12 | 27.2 | — | 25.3 | 95.4 | |
| ViECap+ToCaEncoder=ViT-B/32, Decoder=GPT2-base2025.12 | 27.1 | — | 25.4 | 95 | |
| CapDecEncoder=ResNet50, Decoder=GPT2-large2025.12 | 26.4 | 69.2 | 25.1 | 91.8 | |
| WS-ClipCapEncoder=N/A, Decoder=GPT22025.12 | 22.1 | 65.5 | 22.2 | 74.6 | |
| LMCapEncoder=ViT-H-14, Decoder=XGLM-2.9B2025.12 | 19.9 | — | 22 | 75.9 | |
| TOMCap (embedding only)Encoder=ViT-L/16, Decoder=GPT2-base2025.12 | 19.1 | 64.1 | 22.1 | 76.6 | |
| MeaCapToTEncoder=ViT-B/32, Decoder=CBART2025.12 | 17.7 | — | 24.3 | 84.8 | |
| MacCapEncoder=ViT-B/32, Decoder=OPT-1.3B2025.12 | 17.4 | 61.4 | 22.3 | 69.7 | |
| CLMEncoder=N/A, Decoder=GPT22025.12 | 15 | 59.3 | 18.7 | 55.7 | |
| MAGICEncoder=ViT-B/32, Decoder=GPT2-small2025.12 | 12.9 | 56.8 | 17.4 | 49.3 | |
| MeaCapTFEncoder=ViT-B/32, Decoder=CBART2025.12 | 9.1 | — | 20.6 | 56.9 | |
| DeCapEncoder=ViT-B/32, Decoder=Transformer2025.12 | 8.9 | — | 17.5 | 50.6 | |
| TOMCap (no training, retrieval)Encoder=ViT-L/16, Decoder=GPT2-base2025.12 | 8.6 | 30.4 | 8.4 | 15.2 | |
| ZeroCapEncoder=ViT-B/32, Decoder=GPT2-medium2025.12 | 7 | 49.8 | 15.4 | 34.5 | |
| SocraticEncoder=ViT-L/14, Decoder=GPT-32025.12 | 6.9 | — | 15 | 44.5 | |
| ConZicEncoder=ViT-B/32, Decoder=BERT-base2025.12 | 1.3 | — | 11.2 | 13.3 |