Text-to-Video retrieval on LSMDC (test)
2,690R@5CE
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| CEProtocol=Fine-Tuned, Note=use extra labeled data in the form of pre-trained semantic embeddings2020.03 | 2,690 | 1,120 | 3,480 | 25 | — | — | |
| MEE + COCO + Facetraining_augmentation=COCO, additional_descriptors=Face2018.04 | 2,560 | 1,010 | 3,460 | 27 | — | — | |
| MoEEProtocol=Fine-Tuned2020.03 | 2,560 | 1,010 | 3,460 | 27 | — | — | |
| CCA (FV HGLMM)features=same features2018.04 | 2,170 | 750 | 3,100 | 33 | — | — | |
| JSFusionNote=LSMDC17 Winner2018.04 | 2,120 | 910 | 3,410 | 36 | — | — | |
| JSFusionProtocol=Fine-Tuned2020.03 | 2,120 | 910 | 3,410 | 36 | — | — | |
| Soft Max MarginProtocol=Fine-Tuned2020.03 | 1,980 | 640 | 2,840 | 39 | — | — | |
| HTM-PT*Protocol=Fine-Tuned2020.03 | 1,960 | 710 | 2,790 | 40 | — | — | |
| Miech et al.2018.04 | 1,920 | 730 | 2,710 | 52 | — | — | |
| HTM-no-PTProtocol=Fine-Tuned2020.03 | 1,880 | 580 | 2,840 | 45 | — | — | |
| CT-SAN2018.04 | 1,630 | 510 | 2,520 | 46 | — | — | |
| SNUVLNote=LSMDC16 Winner2018.04 | 1,470 | 360 | 2,390 | 50 | — | — | |
| C+LSTM+SA+FC72018.04 | 1,300 | 420 | 1,950 | 90 | — | — | |
| Soft Max MarginProtocol=Zero-Shot2020.03 | 1,160 | 420 | 1,710 | 119 | — | — | |
| HTM-PT*Protocol=Zero-Shot2020.03 | 980 | 400 | 1,400 | 137 | — | — | |
| UMT-LExample=425M2023.06 | 65.5 | 43 | 73 | — | — | — | |
| COSAExample=415M2023.06 | 60.4 | 39.4 | 67.7 | — | — | — | |
| COSA-LExample=417M2023.06 | 59.1 | 38.6 | 67.4 | — | — | — | |
| VALOR_L#Example=33.5M, Mod=V+A, Dual Softmax (+DSL)=true2023.04 | 56 | 34.2 | 64.1 | — | — | — | |
| mPLUG-2Example=417M2023.06 | 55.2 | 34.4 | 65.1 | — | — | — | |
| All-in-oneExample=138M2023.06 | 53.7 | 22.4 | 67.7 | — | — | — | |
| VALOR-LExample=433.5M2023.06 | 52.8 | 31.8 | 62.4 | — | — | — | |
| VALOR_L#Example=33.5M, Mod=V+A, Dual Softmax (+DSL)=false2023.04 | 52.8 | 31.8 | 62.4 | — | — | — | |
| CLIP-ViPBackbone=ViT-B/16, Post-processing DSL=true2022.09 | 51.4 | 30.7 | 60.6 | — | 5 | — | |
| CLIP-VIP#Example=100M, Mod=V, Dual Softmax (+DSL)=true2023.04 | 51.4 | 30.7 | 60.6 | — | — | — | |
| COSA-BExample=17M2023.06 | 50.9 | 31.2 | 57.8 | — | — | — | |
| CLIP-ViPBackbone=ViT-B/16, Post-processing DSL=false2022.09 | 50.6 | 29.4 | 59 | — | 5 | — | |
| CLIP-VIPExample=500M2023.06 | 50.6 | 29.4 | 59 | — | — | — | |
| HiTeA# PT Data=17M, Inference Setting=Fine-tuning2022.12 | 50.3 | 28.7 | 59 | — | — | — | |
| HiTeAExample=17M2023.06 | 50.3 | 28.7 | 59 | — | — | — | |
| Random baseline2018.04 | 50 | 10 | 100 | 500 | — | — | |
| RandomProtocol=Zero-Shot2020.03 | 50 | 10 | 100 | 500 | — | — | |
| X-CLIPExample=400M2023.06 | 48.4 | 26.1 | 46.7 | — | — | — | |
| X-CLIPMod=V2023.04 | 48.4 | 26.1 | 46.7 | — | — | — | |
| DRLBackbone=ViT-B/162022.09 | 47.6 | 26.5 | 56.8 | — | 7 | — | |
| DCRMod=V2023.04 | 47.6 | 26.5 | 56.8 | — | — | — | |
| COSA-BExample=5M2023.06 | 46.7 | 27.3 | 55.2 | — | — | — | |
| MDMMT-22022.03 | 46.7 | 26.9 | 55.9 | 6.7 | 48 | — | |
| VAST, HowToCaption-finetunedV. Encoder=ViT-G2023.10 | 46.5 | 27.7 | 54.6 | 7 | — | — | |
| CLIP-ViPBackbone=ViT-B/32, Post-processing DSL=true2022.09 | 46.4 | 26 | 54.9 | — | 8 | — | |
| LAVENDERExample=30M2023.06 | 46.4 | 26.1 | 57.3 | — | — | — | |
| LAVENDER#Example=30M, Mod=V2023.04 | 46.4 | 26.1 | 57.3 | — | — | — | |
| HunYuan_tvrMod=V, Dual Softmax (+DSL)=true2023.04 | 46.4 | 29.7 | 55.4 | — | — | — | |
| CenterCLIP (k-medoids++, Bt = 4, 160)Backbone=ViT-B/16, MeM. GB=17.6, speed ms=86.52022.05 | 46.2 | 24.2 | 55.9 | 8 | 47.3 | — | |
| HiTeA# PT Data=5M, Inference Setting=Fine-tuning2022.12 | 46.2 | 27.1 | 54.5 | — | — | — | |
| CenterCLIPBackbone=ViT-B/162022.09 | 46.2 | 24.2 | 55.9 | — | 8 | — | |
| CAMOEBackbone=ViT-B/32, Post-processing DSL=true2022.09 | 46.1 | 25.9 | 53.7 | — | — | — | |
| CAMoE + DSL2023.07 | 46.1 | 25.9 | 53.7 | — | 54.4 | — | |
| TEFAL2023.07 | 46.1 | 26.8 | 56.5 | 7 | 44.4 | — | |
| CAMoE2022.03 | 46.1 | 25.9 | 53.7 | — | 54.4 | — | |
| VALOR_B#Example=6.5M, Mod=V+A2023.04 | 45.8 | 25.1 | 55.2 | — | — | — | |
| DRLBackbone=ViT-B/322022.09 | 45.7 | 24.9 | 55.3 | — | 7 | — | |
| CLIP-ViPBackbone=ViT-B/32, Post-processing DSL=false2022.09 | 45.3 | 25.6 | 54.4 | — | 8 | — | |
| CLIP4clip (meanP)Backbone=ViT-B/16, MeM. GB=25.7, speed ms=59.62022.05 | 45 | 24.1 | 55.1 | 8 | 51.1 | — | |
| CLIP4ClipBackbone=ViT-B/162022.09 | 45 | 24.1 | 55.1 | — | 8 | — | |
| LSDOYear=20252026.06 | 44.8 | 26.4 | 55.1 | 8 | 52.1 | — | |
| DREAM2026.06 | 44.6 | 27.3 | 56.4 | 8 | 49.6 | — | |
| CloverPre-training dataset=W2M+C3M, Evaluation Protocol=Fine-tune2022.07 | 44 | 24.8 | 54.5 | 8 | — | — | |
| LAVENDER# PT Data=5M, Inference Setting=Fine-tuning2022.12 | 43.8 | 22.2 | 53.5 | — | — | — | |
| mPLUG-2Zero-shot=true, Pre-training Data=17M2023.02 | 43.8 | 24.1 | 52 | — | — | — | |
| mPLUG-2V. Encoder=ViT-L2023.10 | 43.8 | 24.1 | 52 | — | — | — | |
| X-Pool2023.07 | 43.7 | 25.2 | 53.5 | 8 | 53.2 | — | |
| VIOLETv2Example=5M2023.06 | 43.5 | 24 | 54.1 | — | — | — | |
| X-CLIP# PT Data=400M, Inference Setting=Fine-tuning2022.12 | 43 | 23.3 | — | — | — | — | |
| Unmasked Teacher-17MV. Encoder=ViT-L2023.10 | 43 | 25.2 | 50.5 | — | — | — | |
| CAMOE2022.07 | 42.6 | 22.5 | 50.9 | — | — | 116 | |
| CAMoEinverted softmax=false2022.07 | 42.6 | 22.5 | 50.9 | — | 56.5 | — | |
| XPoolBackbone=ViT-B/322022.09 | 42.6 | 22.7 | 51.2 | — | 10 | — | |
| TS2-Net2022.07 | 42.3 | 23.4 | 50.9 | 9 | — | 116.6 | |
| TS2-Netinverted softmax=false2022.07 | 42.3 | 23.4 | 50.9 | 9 | 56.9 | — | |
| TS2-Net2023.07 | 42.3 | 23.4 | 50.9 | 9 | 56.9 | — | |
| TS2-NetMod=V2023.04 | 42.3 | 23.4 | 50.9 | — | — | — | |
| CloverExample=5M2023.06 | 42 | 22.7 | 52.6 | — | — | — | |
| CLIP4Clip-seqLSTMTrainD=WIT + LSMDC, E2E=true2021.04 | 41.8 | 21.6 | 49.8 | 11 | 58 | — | |
| Clip4Clip# PT Data=400M, Inference Setting=Fine-tuning2022.12 | 41.8 | 21.6 | 49.8 | — | — | — | |
| CLIP4ClipBackbone=ViT-B/322022.09 | 41.8 | 21.6 | 49.8 | — | 11 | — | |
| CLIP4Clip2022.03 | 41.8 | 21.6 | 49.8 | — | 58 | — | |
| SRL-CLIPBackbone=ViT-L/14, DSize=23k, Params=400M, Zero-shot=true, Pretraining Dataset=VidSitu2024.01 | 41.7 | — | 48.7 | 12 | — | — | |
| Align and TellYear=20232026.06 | 41.2 | 23.1 | 49.6 | 11 | — | — | |
| BridgeFormerEvaluation Protocol=Fine-tuned2022.01 | 41.1 | 21.8 | 50.6 | 10 | 60.5 | — | |
| CenterCLIP (k-medoids++, Bt = 6, 49)Backbone=ViT-B/32, MeM. GB=16.4, speed ms=23.92022.05 | 41.1 | 21.9 | 50.7 | 10 | 55.6 | — | |
| CenterCLIP2023.07 | 41.1 | 21.9 | 50.7 | 10 | 55.6 | — | |
| UMPYear=20242026.06 | 41.1 | 21.6 | 49.9 | — | 59.5 | — | |
| CLIP4Clip-seqTransfTrainD=WIT + LSMDC, E2E=true2021.04 | 41 | 22.6 | 49.1 | 11 | 61 | — | |
| CLIP4clip (seqTransf)2022.05 | 41 | 22.6 | 49.1 | 11 | 61 | — | |
| CLIP4Clipinverted softmax=false2022.07 | 41 | 22.6 | 49.1 | 11 | 61 | — | |
| CLIP4ClipExample=400M2023.06 | 41 | 22.6 | 49.1 | — | — | — | |
| CLIP4ClipMod=V2023.04 | 41 | 22.6 | 49.1 | — | — | — | |
| Clip4ClipYear=20222026.06 | 41 | 22.6 | 49.1 | 11 | 61 | — | |
| CenterCLIP (spectral, Bt = 6, 49)Backbone=ViT-B/32, MeM. GB=16.4, speed ms=40.82022.05 | 40.9 | 21.6 | 49.3 | 11 | 57.2 | — | |
| VAST†V. Encoder=ViT-G2023.10 | 40.9 | 23.2 | 48.9 | 12 | — | — | |
| CLIP4clip (meanP)Backbone=ViT-B/32, MeM. GB=20.8, speed ms=24.42022.05 | 40.2 | 20.1 | 48.4 | 12 | 57.1 | — | |
| QB-Norminverted softmax=true2022.07 | 40.1 | 22.4 | 49.5 | 11 | — | — | |
| QB-Norm+CLIP4Clip2022.03 | 40.1 | 22.4 | 49.5 | 11 | — | — | |
| CenterCLIP (k-medoids++, Bt = 4, 49)Backbone=ViT-B/32, MeM. GB=15, speed ms=22.92022.05 | 39.8 | 21.7 | 49.8 | 11 | 54.8 | — | |
| CenterCLIPBackbone=ViT-B/322022.09 | 39.8 | 21.7 | 49.8 | — | 11 | — | |
| CenterCLIP (spectral, Bt = 4, 49)Backbone=ViT-B/32, MeM. GB=15, speed ms=43.62022.05 | 39.7 | 21.4 | 49.4 | 11 | 55.9 | — | |
| VALOR_E#Example=5.5M, Mod=V2023.04 | 39.1 | 20 | 49 | — | — | — | |
| CLIP4Clip-meanPTrainD=WIT + LSMDC, E2E=true2021.04 | 38.9 | 20.7 | 47.2 | 13 | 65.3 | — | |
| CLIP4ClipEvaluation Protocol=Fine-tuned2022.01 | 38.9 | 20.7 | 47.2 | 13 | 65.3 | — |