Video Captioning on MSR-VTT (test)
104.2CIDErComHeat
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| ComHeat2024.09 | 104.2 | 63.9 | 42.7 | — | — | 80.3 | — | — | |
| GL-RG+IT2024.09 | 101 | 60.5 | 38.9 | — | — | 76.4 | — | — | |
| O2NA2024.09 | 96.4 | 55.4 | 37.4 | — | — | 74.5 | — | — | |
| ORG-TRL2024.09 | 95.2 | 54.3 | 36.4 | — | — | 73.9 | — | — | |
| POSRL2024.09 | 91 | 53.9 | 34.9 | — | — | 72.1 | — | — | |
| OA-BTG2024.09 | 90.6 | 56.9 | 36.2 | — | — | — | — | — | |
| mPLUG2+Quantization=None (32-bit weights), Text Decoder=LLaMa2025.12 | 83.1 | 62.4 | 70.8 | 36.8 | — | — | — | — | |
| mPLUG2+Quantization=8-bit weights, Text Decoder=LLaMa2025.12 | 82.5 | 60.7 | 68.6 | 35.9 | — | — | — | — | |
| mPLUG2Quantization=None (32-bit weights), Text Decoder=BERT2025.12 | 79.4 | 56.9 | 68.2 | 34 | — | — | — | — | |
| SOTAEvaluation Protocol=Pre-training & finetuning2022.09 | 75.9 | — | — | — | — | — | — | — | |
| mPLUG2Quantization=8-bit weights, Text Decoder=BERT2025.12 | 75.3 | 52.4 | 65.4 | 30.8 | — | — | — | — | |
| EVLGenPre-training=Video, Training Strategy=SCST2023.10 | 74 | 49.2 | 33 | 66.5 | — | — | — | — | |
| EVLGenPre-training=Video2023.10 | 69.8 | 48.3 | 32.6 | 65.8 | — | — | — | — | |
| EVLGenPre-training=Image2023.10 | 68.4 | 47.6 | 32.4 | 65.3 | — | — | — | — | |
| BaselinePooling=mean2023.10 | 67.8 | 47.3 | 32.2 | 65 | — | — | — | — | |
| BaselinePooling=concat2023.10 | 65.5 | 44.4 | 31.9 | 64.1 | — | — | — | — | |
| Vid2SeqTrained Parameters=314M, Pretraining Data=YT-Temporal-1B2023.02 | 64.6 | — | 30.8 | — | — | — | — | — | |
| VideoCoCaVersion=open2023.10 | 63 | 48.5 | 31.4 | 64.8 | — | — | — | — | |
| Vid2SeqTrained Parameters=314M, Pretraining Data=HowTo100M2023.02 | 61.5 | — | 30.4 | — | — | — | — | — | |
| MV-GPTPre-trained parts=Encoder+Decoder, Inputs=Video+Text2022.01 | 60 | 48.92 | 38.66 | 64 | — | — | — | — | |
| MV-GPTTrained Parameters=354M, Pretraining Data=HowTo100M2023.02 | 60 | — | 29.9 | — | — | — | — | — | |
| Video-LLaMA2023.10 | 59.3 | 47.7 | 29.6 | 63.7 | — | — | — | — | |
| CLIP-DCDEvaluation Mode=Ours2021.11 | 58.7 | 48.2 | 31.3 | 64.8 | — | — | — | — | |
| VideoChat2023.10 | 58 | 46.5 | 29.5 | 63.4 | — | — | — | — | |
| Vid2SeqTrained Parameters=314M, Pretraining Data=Ø2023.02 | 57.2 | — | 30 | — | — | — | — | — | |
| CLIP-BaseEvaluation Mode=Ours2021.11 | 57 | 47.3 | 30.9 | 64 | — | — | — | — | |
| audiovisual dual stream retrieval modelEvaluation Protocol=Finetuned, Pre-training Data=VideoCC12M, Modality=V2022.04 | 56 | 47.21 | 37.7 | — | — | — | — | — | |
| OursPretraining Data=HowTo100M, Modality=V, Evaluation Protocol=Finetuned2022.04 | 55 | 47.33 | 37.11 | — | — | — | — | — | |
| OursPretraining Data=VideoCC3M, Modality=V, Evaluation Protocol=Finetuned2022.04 | 55 | 45.47 | 36.96 | — | — | — | — | — | |
| audiovisual dual stream retrieval modelEvaluation Protocol=Finetuned, Pre-training Data=VideoCC3M, Modality=V2022.04 | 55 | 45.47 | 36.96 | — | — | — | — | — | |
| audiovisual dual stream retrieval modelEvaluation Protocol=Finetuned, Pre-training Data=HowTo100M, Modality=V2022.04 | 55 | 47.33 | 37.11 | — | — | — | — | — | |
| Base model_capBackbone=CLIP [53], modality=image only2022.11 | 54.6 | 45.2 | 29.8 | 63 | — | — | — | — | |
| EMCLBackbone=CLIP [53], modality=image only2022.11 | 54.6 | 45.3 | 30.2 | 63.2 | — | — | — | — | |
| SWINBERTframes=642021.11 | 53.8 | — | — | — | — | — | — | — | |
| SWINBERT2D Appearance=VidSwin2021.11 | 53.8 | 41.9 | 29.9 | 62.1 | — | — | — | — | |
| SwinBERTTrained Parameters=229M, Pretraining Data=Ø2023.02 | 53.8 | — | 29.9 | — | — | — | — | — | |
| VNS-GRUModality=V, Evaluation Protocol=Finetuned2022.04 | 53 | 45.3 | 29.9 | — | — | — | — | — | |
| VNS-GRUEvaluation Protocol=Finetuned, Pre-training Data=None, Modality=V2022.04 | 53 | 45.3 | 29.9 | — | — | — | — | — | |
| VNS-GRUPre-trained parts=None, Inputs=Video2022.01 | 53 | 45.3 | 29.9 | 63.4 | — | — | — | — | |
| SOTA2021.11 | 52.9 | — | — | — | — | — | — | — | |
| OpenBook2D Appearance=IncepResnetV2, 3D Motion=C3D2021.11 | 52.9 | 42.8 | 29.3 | 61.7 | — | — | — | — | |
| OpenBookYear=20212021.11 | 52.9 | 42.8 | 29.3 | 61.7 | — | — | — | — | |
| OpenBookmodality=image only2022.11 | 52.9 | 42.8 | 29.3 | 61.7 | — | — | — | — | |
| DCDmodality=image only2022.11 | 52.8 | 43.4 | 29.6 | 61.8 | — | — | — | — | |
| TextKGEvaluation Mode=micro-level2023.03 | 52.4 | 43.7 | 29.6 | — | — | 62.4 | — | — | |
| APMLYear=20212021.11 | 52.2 | 43.8 | 30.3 | 63.6 | — | — | — | — | |
| DECEMBERTPretraining Data=HowTo100M, Modality=V, Evaluation Protocol=Finetuned2022.04 | 52 | 45.2 | 29.7 | — | — | — | — | — | |
| DECEMBERTEvaluation Protocol=Finetuned, Pre-training Data=HowTo100M, Modality=V2022.04 | 52 | 45.2 | 29.7 | — | — | — | — | — | |
| DECEMBERTPre-trained parts=Encoder, Inputs=Video2022.01 | 52 | 45.2 | 29.7 | 64.7 | — | — | — | — | |
| HMNEvaluation Mode=micro-level2023.03 | 51.5 | 43.5 | 29 | — | — | 62.7 | — | — | |
| MGCMPYear=20212021.11 | 51.4 | 41.7 | 28.9 | 62.1 | — | — | — | — | |
| NACFYear=20212021.11 | 51.4 | 42 | 28.7 | — | — | — | — | — | |
| MGCMPmodality=image only2022.11 | 51.4 | 41.7 | 28.9 | 62.1 | — | — | — | — | |
| MGRMPEvaluation Mode=micro-level2023.03 | 51.4 | 41.7 | 28.9 | — | — | 62.1 | — | — | |
| ARB-ACLYear=20222021.11 | 51.3 | 42.6 | 28.9 | 61.5 | — | — | — | — | |
| ARB-ACLmodality=image only2022.11 | 51.3 | 42.6 | 28.9 | 61.5 | — | — | — | — | |
| SAM-SSModality=V, Evaluation Protocol=Finetuned2022.04 | 51 | 43.8 | 28.9 | — | — | — | — | — | |
| ORG-TRLModality=V, Evaluation Protocol=Finetuned2022.04 | 51 | 43.6 | 28.8 | — | — | — | — | — | |
| ORG-TRLEvaluation Protocol=Finetuned, Pre-training Data=None, Modality=V2022.04 | 51 | 43.6 | 28.8 | — | — | — | — | — | |
| SAM-SSPre-trained parts=None, Inputs=Video2022.01 | 51 | 43.8 | 28.9 | 62.4 | — | — | — | — | |
| ORG-TRLPre-trained parts=None, Inputs=Video2022.01 | 51 | 43.6 | 28.8 | 62.8 | — | — | — | — | |
| SAAT2D Appearance=IncepResnetV2, 3D Motion=C3D2021.11 | 51 | 39.9 | 27.7 | 61.2 | — | — | — | — | |
| ORG-TRLAppearance=InceptionResnetV2, Motion=C3D, Object=FasterRCNN2020.02 | 50.9 | 43.6 | 28.8 | 62.1 | — | — | — | — | |
| ORG-TRL2D Appearance=IncepResnetV2, 3D Motion=C3D, Object Detection=FasterRCNN2021.11 | 50.9 | 43.6 | 28.8 | 62.1 | — | — | — | — | |
| ORG-TRLYear=20202021.11 | 50.9 | 43.6 | 28.8 | 62.1 | — | — | — | — | |
| ORG-TRLmodality=image only2022.11 | 50.9 | 43.6 | 28.8 | 62.1 | — | — | — | — | |
| ORG-TRLEvaluation Mode=micro-level2023.03 | 50.9 | 43.6 | 28.8 | — | — | 62.1 | — | — | |
| UniVL#VideoPT=1.2M, ASR=No, Evaluation Protocol=Fine-tuning2022.05 | 50.1 | 42 | 29 | 61 | 70.2 | — | — | — | |
| UniVLPretraining Data=HowTo100M, Modality=V+T, Evaluation Protocol=Finetuned2022.04 | 50 | 41.79 | 28.94 | — | — | — | — | — | |
| UniVLEvaluation Protocol=Finetuned, Pre-training Data=HowTo100M, Modality=V+T2022.04 | 50 | 41.79 | 28.94 | — | — | — | — | — | |
| UniVLPre-trained parts=Encoder+Decoder, Inputs=Video+Text2022.01 | 50 | 41.79 | 28.94 | 60.78 | — | — | — | — | |
| SGNEvaluation Mode=micro-level2023.03 | 49.5 | 40.8 | 28.3 | — | — | 60.8 | — | — | |
| PMI-CAP2D Appearance=IncepResnetV2, 3D Motion=C3D2021.11 | 49.4 | 42.1 | 28.7 | — | — | — | — | — | |
| OracleVLM=SmolVLM2-2.2B, k (frame budget)=2, Zero-shot=true2026.05 | 49.35 | 37.44 | 27.81 | 59.58 | — | — | — | — | |
| CSTAVLM=Qwen2.5-VL-7B, k (frame budget)=4, Zero-shot=true2026.05 | 49.12 | 37.31 | 27.41 | 58.43 | — | — | — | — | |
| POS+VCTYear=2019, Appearance=InceptionResnetV2, Motion=C3D2020.02 | 49.1 | 42.3 | 29.7 | 62.8 | — | — | — | — | |
| POS+VCT2D Appearance=IncepResnetV2, 3D Motion=C3D2021.11 | 49.1 | 42.3 | 29.7 | 62.8 | — | — | — | — | |
| POS+CGModality=V, Evaluation Protocol=Finetuned2022.04 | 49 | 42 | 28.2 | — | — | — | — | — | |
| POS+VCTModality=V, Evaluation Protocol=Finetuned2022.04 | 49 | 42.3 | 29.7 | — | — | — | — | — | |
| POS+CGPre-trained parts=None, Inputs=Video2022.01 | 49 | 42 | 28.2 | 61.6 | — | — | — | — | |
| POS+VCTPre-trained parts=None, Inputs=Video2022.01 | 49 | 42.3 | 29.7 | 62.8 | — | — | — | — | |
| OracleVLM=SmolVLM2-2.2B, k (frame budget)=4, Zero-shot=true2026.05 | 48.99 | 37.66 | 27.9 | 59.74 | — | — | — | — | |
| POS+CGYear=2019, Appearance=InceptionResnetV2, Motion=OpticalFlow2020.02 | 48.7 | 42 | 28.2 | 61.6 | — | — | — | — | |
| POS+CG2D Appearance=IncepResnetV2, 3D Motion=OpticalFlow2021.11 | 48.7 | 42 | 28.2 | 61.6 | — | — | — | — | |
| POS-CGEvaluation Mode=micro-level2023.03 | 48.7 | 42 | 28.2 | — | — | 61.6 | — | — | |
| OracleVLM=SmolVLM2-2.2B, k (frame budget)=1, Zero-shot=true2026.05 | 48.33 | 37.19 | 27.65 | 59.47 | — | — | — | — | |
| GRU-EVEYear=2019, Appearance=InceptionResnetV2, Motion=C3D, Object=YOLO2020.02 | 48.1 | 38.3 | 28.4 | 60.7 | — | — | — | — | |
| GRU-EVE2D Appearance=IncepResnetV2, 3D Motion=C3D, Object Detection=YOLO2021.11 | 48.1 | 38.3 | 28.4 | 60.7 | — | — | — | — | |
| MGSAPre-trained parts=None, Inputs=Video2022.01 | 48 | 42.4 | 27.6 | — | — | — | — | — | |
| SibNetYear=2019, Appearance=GoogleNet2020.02 | 47.5 | 40.9 | 27.5 | 60.2 | — | — | — | — | |
| MGSAYear=2019, Appearance=InceptionResnetV2, Motion=C3D2020.02 | 47.5 | 42.4 | 27.6 | — | — | — | — | — | |
| SibNet2D Appearance=GoogleNet2021.11 | 47.5 | 40.9 | 27.5 | 60.2 | — | — | — | — | |
| MGSA2D Appearance=IncepResnetV2, 3D Motion=C3D2021.11 | 47.5 | 42.4 | 27.6 | — | — | — | — | — | |
| MGSAEvaluation Mode=micro-level2023.03 | 47.5 | 42.4 | 27.6 | — | — | — | — | — | |
| MARNYear=2019, Appearance=ResNet-101, Motion=C3D2020.02 | 47.1 | 40.4 | 28.1 | 60.7 | — | — | — | — | |
| STG-KD2D Appearance=ResNet101, 3D Motion=I3D, Object Detection=FasterRCNN2021.11 | 47.1 | 40.5 | 28.3 | 60.9 | — | — | — | — | |
| STG-KDmodality=image only2022.11 | 47.1 | 40.5 | 28.3 | 60.9 | — | — | — | — | |
| STG-KDEvaluation Mode=micro-level2023.03 | 47.1 | 40.5 | 28.3 | — | — | 60.9 | — | — | |
| OA-BTGPre-trained parts=None, Inputs=Video2022.01 | 47 | 41.4 | 28.2 | — | — | — | — | — | |
| OA-BTGYear=2019, Appearance=ResNet-200, Object=Mask-RCNN2020.02 | 46.9 | 41.4 | 28.2 | — | — | — | — | — |