Image Captioning on NoCaps (test)
124.8CIDErGIT2
Evaluation Results
| Method | Links | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| GIT2Number of parameters=5.1B2022.09 | 124.8 | — | — | — | — | — | — | — | — | — | — | |
| GIT2Parameters=5.1B, Fine-tuning=per-task, OCR_pipeline=without2023.05 | 124.8 | — | — | — | — | — | — | — | — | — | — | |
| GIT2Number of parameters=5.1B2023.05 | 124.8 | — | — | — | — | — | — | — | — | — | — | |
| PaLINumber of parameters=17B2022.09 | 124.4 | — | — | — | — | — | — | — | — | — | — | |
| PaLIParameters=17B, Fine-tuning=per-task, OCR_pipeline=without2023.05 | 124.4 | — | — | — | — | — | — | — | — | — | — | |
| PaLINumber of parameters=17B2023.05 | 124.4 | — | — | — | — | — | — | — | — | — | — | |
| PaLI-XParameters=55B, Fine-tuning=per-task, OCR_pipeline=without2023.05 | 124.3 | — | — | — | — | — | — | — | — | — | — | |
| PaLI-XNumber of parameters=55B2023.05 | 124.3 | — | — | — | — | — | — | — | — | — | — | |
| GITNumber of parameters=0.7B2022.09 | 123.4 | — | — | — | — | — | — | — | — | — | — | |
| GITNumber of parameters=0.7B2023.05 | 123.4 | — | — | — | — | — | — | — | — | — | — | |
| CoCaNumber of parameters=2.1B2022.09 | 120.6 | — | — | — | — | — | — | — | — | — | — | |
| CoCaNumber of parameters=2.1B2023.05 | 120.6 | — | — | — | — | — | — | — | — | — | — | |
| LEMONNumber of parameters=0.7B2022.09 | 114.3 | — | — | — | — | — | — | — | — | — | — | |
| SimVLM2022.09 | 110.3 | — | — | — | — | — | — | — | — | — | — | |
| SimVLM2023.05 | 110.3 | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-1.5-7BTokens Retained=576, Retention Ratio=100%2026.06 | 105.61 | — | — | — | — | — | — | — | — | — | — | |
| ALVTSTokens Retained=192, Reduction Ratio=67%2026.06 | 104.5 | — | — | — | — | — | — | — | — | — | — | |
| DARTTokens Retained=192, Reduction Ratio=67%2026.06 | 103.84 | — | — | — | — | — | — | — | — | — | — | |
| PDropTokens Retained=192, Reduction Ratio=67%2026.06 | 103.76 | — | — | — | — | — | — | — | — | — | — | |
| DARTTokens Retained=128, Reduction Ratio=78%2026.06 | 102.86 | — | — | — | — | — | — | — | — | — | — | |
| ALVTSTokens Retained=128, Reduction Ratio=78%2026.06 | 102.3 | — | — | — | — | — | — | — | — | — | — | |
| FastVTokens Retained=192, Reduction Ratio=67%2026.06 | 102.26 | — | — | — | — | — | — | — | — | — | — | |
| DARTTokens Retained=64, Reduction Ratio=89%2026.06 | 97.7 | — | — | — | — | — | — | — | — | — | — | |
| ALVTSTokens Retained=64, Reduction Ratio=89%2026.06 | 97.6 | — | — | — | — | — | — | — | — | — | — | |
| FastVTokens Retained=128, Reduction Ratio=78%2026.06 | 97.06 | — | — | — | — | — | — | — | — | — | — | |
| PDropTokens Retained=128, Reduction Ratio=78%2026.06 | 97.01 | — | — | — | — | — | — | — | — | — | — | |
| PDropTokens Retained=64, Reduction Ratio=89%2026.06 | 85.45 | — | — | — | — | — | — | — | — | — | — | |
| FastVTokens Retained=64, Reduction Ratio=89%2026.06 | 79.01 | — | — | — | — | — | — | — | — | — | — | |
| VISEModel Scale=Qwen3-VL-32B-Instruct2026.06 | 35.95 | — | — | — | — | — | — | — | — | 34.41 | 50.74 | |
| VisionZero-RWModel Scale=Qwen3-VL-32B-Instruct2026.06 | 35.12 | — | — | — | — | — | — | — | — | 34.29 | 50.69 | |
| VISEModel Scale=Qwen3-VL-8B-Instruct2026.06 | 34.98 | — | — | — | — | — | — | — | — | 30.31 | 48.66 | |
| VISEModel Scale=Qwen3-VL-4B-Instruct2026.06 | 34.97 | — | — | — | — | — | — | — | — | 29.6 | 48.86 | |
| VISEModel Scale=Qwen3-VL-2B-Instruct2026.06 | 34.25 | — | — | — | — | — | — | — | — | 28.6 | 48.69 | |
| VisionZero-RWModel Scale=Qwen3-VL-8B-Instruct2026.06 | 33.87 | — | — | — | — | — | — | — | — | 30.27 | 48.59 | |
| VisionZero (Chart)Model Scale=Qwen3-VL-32B-Instruct2026.06 | 33.14 | — | — | — | — | — | — | — | — | 34.1 | 50.66 | |
| VisionZero (CLEVR)Model Scale=Qwen3-VL-32B-Instruct2026.06 | 32.41 | — | — | — | — | — | — | — | — | 34.02 | 50.61 | |
| iReasonerModel Scale=Qwen3-VL-32B-Instruct2026.06 | 31.68 | — | — | — | — | — | — | — | — | 33.97 | 50.53 | |
| VisionZero (Chart)Model Scale=Qwen3-VL-8B-Instruct2026.06 | 31.47 | — | — | — | — | — | — | — | — | 30.22 | 48.53 | |
| VisionZero (CLEVR)Model Scale=Qwen3-VL-8B-Instruct2026.06 | 29.92 | — | — | — | — | — | — | — | — | 30.14 | 48.44 | |
| VisPlayModel Scale=Qwen3-VL-32B-Instruct2026.06 | 28.62 | — | — | — | — | — | — | — | — | 33.51 | 50.09 | |
| VisionZero-RWModel Scale=Qwen3-VL-4B-Instruct2026.06 | 28.61 | — | — | — | — | — | — | — | — | 29.15 | 46.09 | |
| iReasonerModel Scale=Qwen3-VL-8B-Instruct2026.06 | 28.44 | — | — | — | — | — | — | — | — | 30.09 | 48.21 | |
| VisionZero (Chart)Model Scale=Qwen3-VL-4B-Instruct2026.06 | 28.12 | — | — | — | — | — | — | — | — | 29.22 | 46.13 | |
| EvoLMMModel Scale=Qwen3-VL-32B-Instruct2026.06 | 28.03 | — | — | — | — | — | — | — | — | 33.44 | 50.28 | |
| BaseModel Scale=Qwen3-VL-32B-Instruct2026.06 | 27.74 | — | — | — | — | — | — | — | — | 33.38 | 50.02 | |
| VisionZero (CLEVR)Model Scale=Qwen3-VL-4B-Instruct2026.06 | 27.05 | — | — | — | — | — | — | — | — | 28.71 | 45.74 | |
| VTWTokens Retained=192, Reduction Ratio=67%2026.06 | 26.66 | — | — | — | — | — | — | — | — | — | — | |
| VisPlayModel Scale=Qwen3-VL-8B-Instruct2026.06 | 26.41 | — | — | — | — | — | — | — | — | 29.98 | 47.96 | |
| iReasonerModel Scale=Qwen3-VL-4B-Instruct2026.06 | 25.52 | — | — | — | — | — | — | — | — | 28.73 | 45.78 | |
| VisPlayModel Scale=Qwen3-VL-4B-Instruct2026.06 | 25.46 | — | — | — | — | — | — | — | — | 29.1 | 45.72 | |
| EvoLMMModel Scale=Qwen3-VL-4B-Instruct2026.06 | 25.36 | — | — | — | — | — | — | — | — | 28.61 | 45.73 | |
| EvoLMMModel Scale=Qwen3-VL-8B-Instruct2026.06 | 24.89 | — | — | — | — | — | — | — | — | 29.95 | 48.31 | |
| BaseModel Scale=Qwen3-VL-8B-Instruct2026.06 | 24.46 | — | — | — | — | — | — | — | — | 29.91 | 47.82 | |
| VisionZero-RWModel Scale=Qwen3-VL-2B-Instruct2026.06 | 22.61 | — | — | — | — | — | — | — | — | 28.24 | 46.12 | |
| BaseModel Scale=Qwen3-VL-4B-Instruct2026.06 | 22.36 | — | — | — | — | — | — | — | — | 28.91 | 45.95 | |
| VisionZero (Chart)Model Scale=Qwen3-VL-2B-Instruct2026.06 | 20.19 | — | — | — | — | — | — | — | — | 26.74 | 44.33 | |
| VisionZero (CLEVR)Model Scale=Qwen3-VL-2B-Instruct2026.06 | 20.16 | — | — | — | — | — | — | — | — | 28.01 | 44.23 | |
| BaseModel Scale=Qwen3-VL-2B-Instruct2026.06 | 19.52 | — | — | — | — | — | — | — | — | 27.36 | 45.08 | |
| VisPlayModel Scale=Qwen3-VL-2B-Instruct2026.06 | 19.14 | — | — | — | — | — | — | — | — | 27.83 | 45.43 | |
| iReasonerModel Scale=Qwen3-VL-2B-Instruct2026.06 | 18.81 | — | — | — | — | — | — | — | — | 26.75 | 45.96 | |
| EvoLMMModel Scale=Qwen3-VL-2B-Instruct2026.06 | 18.75 | — | — | — | — | — | — | — | — | 26.72 | 45.86 | |
| VTWTokens Retained=64, Reduction Ratio=89%2026.06 | 6.06 | — | — | — | — | — | — | — | — | — | — | |
| VTWTokens Retained=128, Reduction Ratio=78%2026.06 | 5.76 | — | — | — | — | — | — | — | — | — | — | |
| BLIPTrainable Parameters=446M, Transfer=Zero-shot from COCO2025.06 | — | 114.9 | 15.2 | 110.6 | 14.6 | 114.8 | 14.3 | 112.8 | 14.7 | — | — | |
| BLIP-2 ViT FlanT5XLTrainable Parameters=1.1B, LLM=FlanT5-XL, Transfer=Zero-shot from COCO, Speedup=1.00x2025.06 | — | 123.7 | 16.3 | 120.2 | 15.9 | 124.8 | 15.1 | 121.6 | 15.8 | — | — | |
| BLIP-2 ViT OPT2.7BTrainable Parameters=1.1B, LLM=OPT-2.7B, Transfer=Zero-shot from COCO, Speedup=1.07x2025.06 | — | 123 | 15.8 | 117.8 | 15.4 | 123.2 | 15 | 119.6 | 15.4 | — | — | |
| CaMELResult attribution=Results computed by current paper authors2022.09 | — | 88.1 | — | 79.1 | — | 54.6 | — | 75.9 | — | — | — | |
| ClipCapResult attribution=Results computed by current paper authors2022.09 | — | 74.5 | — | 65.6 | — | 47.1 | — | 63.4 | — | — | — | |
| CoCa2022.05 | — | — | — | — | — | — | — | 120.6 | 15.5 | — | — | |
| CoCaTrain Data=4.8B, Protocol=zero-shot2023.11 | — | — | — | — | — | — | — | 120.6 | — | — | — | |
| CogVLMTrain Data=1.5B, Protocol=zero-shot2023.11 | — | — | — | — | — | 128 | — | 126.4 | — | — | — | |
| CogVLMTraining regime=Specialist SOTAS, Backbone=Vicuna-7B2023.11 | — | — | — | — | — | 128 | — | 126.4 | — | — | — | |
| DeeBLIPTrainable Parameters=1.8B, Category=Early Exit models, Speedup=1.41x2025.06 | — | 115.2 | 15.3 | 111.5 | 14.7 | 115.4 | 14.5 | 112.4 | 14.5 | — | — | |
| EVCAPTraining regime=Lightweight-training, Backbone=Vicuna-13B, Memory bank=true2023.11 | — | 114.9 | — | 117 | — | 117.1 | — | 116.8 | — | — | — | |
| FREE ViT FlanT5XLTrainable Parameters=1.4B, LLM=FlanT5-XL, Category=Early Exit models, Speedup=1.51x2025.06 | — | 124.3 | 16.5 | 120 | 15.9 | 125.5 | 15.4 | 122.7 | 16.1 | — | — | |
| FREE ViT OPT2.7BTrainable Parameters=1.5B, LLM=OPT-2.7B, Category=Early Exit models, Speedup=1.63x2025.06 | — | 122.7 | 15.7 | 118.1 | 15.5 | 123.9 | 15.1 | 119.9 | 15.6 | — | — | |
| GITTrain Data=0.8B, Protocol=zero-shot2023.11 | — | — | — | — | — | 122 | — | 123.4 | — | — | — | |
| GIT2Train Data=12.9B, Protocol=zero-shot2023.11 | — | — | — | — | — | 122.3 | — | 124.8 | — | — | — | |
| Human2021.01 | — | 80.6 | 15 | 84.6 | 14.7 | 91.6 | 14.2 | 85.3 | 14.6 | — | — | |
| HumanPre-training data=N/A2021.11 | — | 80.6 | 15 | 84.6 | 14.7 | 91.6 | 14.2 | 85.3 | 14.6 | — | — | |
| Human2022.09 | — | 80.6 | 15 | 84.6 | 14.7 | 91.6 | 14.2 | 85.3 | 14.6 | — | — | |
| HumanProtocol=zero-shot2023.11 | — | — | — | — | — | 91.6 | — | 85.3 | — | — | — | |
| LeeBLIPTrainable Parameters=1.8B, Category=Early Exit models, Speedup=1.38x2025.06 | — | 119.4 | 15.5 | 115.8 | 14.8 | 120.1 | 14.9 | 116.3 | 15.1 | — | — | |
| LEMON2022.05 | — | — | — | — | — | — | — | 114.3 | 14.9 | — | — | |
| LEMONTrain Data=2B, Protocol=zero-shot2023.11 | — | — | — | — | — | 110.1 | — | 114.3 | — | — | — | |
| LEMONhugePre-training data=ALT200M2021.11 | — | 112.8 | 15.2 | 115.5 | 15.1 | 110.1 | 13.7 | 114.3 | 14.9 | — | — | |
| LEMONlargePre-training data=ALT200M2021.11 | — | 111.2 | 15.6 | 112.3 | 15.2 | 105 | 13.6 | 110.9 | 15 | — | — | |
| MuETrainable Parameters=1.8B, Category=Early Exit models, Speedup=1.44x2025.06 | — | 118.1 | 15.4 | 115.3 | 14.8 | 118.7 | 14.8 | 114.8 | 14.9 | — | — | |
| NOC-REKTraining regime=Heavyweight-training, Memory bank=true2023.11 | — | 100 | — | 95.7 | — | 77.4 | — | 93 | — | — | — | |
| OscarModel Scale=Base, SCST optimization=true, Constrained Beam Search=true2022.09 | — | 81.3 | 11.9 | 79.6 | 11.9 | 73.6 | 10.6 | 78.8 | 11.7 | — | — | |
| OscarModel Scale=Large, SCST optimization=true, Constrained Beam Search=true2022.09 | — | 84.8 | 12.1 | 82.1 | 11.5 | 73.8 | 9.7 | 80.9 | 11.3 | — | — | |
| OSCARTraining regime=Heavyweight-training, Memory bank=false2023.11 | — | 81.3 | — | 79.6 | — | 73.6 | — | 78.8 | — | — | — | |
| Oscar + VIVOModel Scale=Base, VIVO pre-training=true, SCST optimization=true, Constrained Beam Search=true2022.09 | — | 89 | 12.9 | 87.8 | 12.6 | 80.1 | 11.1 | 86.6 | 12.4 | — | — | |
| OSCAR_Bsize=base, SCST=true, CBS=true2021.01 | — | 81.3 | 11.9 | 79.6 | 11.9 | 73.6 | 10.6 | 78.8 | 11.7 | — | — | |
| OSCAR_Lsize=large, SCST=true, CBS=true2021.01 | — | 84.8 | 12.1 | 82.1 | 11.5 | 73.8 | 9.7 | 80.9 | 11.3 | — | — | |
| OSCAR-LargeSource=Results copied from publication2022.09 | — | 84.8 | — | 82.1 | — | 73.8 | — | 80.9 | — | — | — | |
| P2CSCST optimization=true, Constrained Beam Search=true2022.09 | — | 101.7 | 15 | 95.7 | 14.4 | 82.5 | 12.2 | 93.5 | 14.1 | — | — | |
| PABEE-BLIPTrainable Parameters=1.8B, Category=Early Exit models, Speedup=1.29x2025.06 | — | 117.7 | 15.4 | 114.2 | 14.8 | 117.6 | 14.7 | 112.9 | 14.6 | — | — | |
| PaLITraining regime=Specialist SOTAS, Backbone=mT5-XXL2023.11 | — | — | — | — | — | — | — | 124.4 | — | — | — | |
| PaLI-17BTrain Data=1.6B, Protocol=zero-shot2023.11 | — | — | — | — | — | — | — | 124.4 | — | — | — |