Video-to-Text Retrieval on VATEX
94.8Recall@1PEAV-L
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| PEAV-LV-Enc Params=0.5B, Video Data=124M, Resolution=336, Frame/ fps=30*, Zero-shot Protocol=true2025.12 | 94.8 | — | — | — | — | |
| PEAV-SV-Enc Params=0.3B, Video Data=124M, Resolution=336, Frame/ fps=30*, Zero-shot Protocol=true2025.12 | 94.5 | — | — | — | — | |
| PEAV-LV-Enc Params=0.5B, Video Data=124M, Resolution=336, Frame/ fps=16, Zero-shot Protocol=true2025.12 | 94.4 | — | — | — | — | |
| PEAV-BV-Enc Params=0.4B, Video Data=124M, Resolution=336, Frame/ fps=30*, Zero-shot Protocol=true2025.12 | 94.4 | — | — | — | — | |
| PEAV-BV-Enc Params=0.4B, Video Data=124M, Resolution=336, Frame/ fps=16, Zero-shot Protocol=true2025.12 | 93.8 | — | — | — | — | |
| PEAV-SV-Enc Params=0.3B, Video Data=124M, Resolution=336, Frame/ fps=16, Zero-shot Protocol=true2025.12 | 93.7 | — | — | — | — | |
| VidVecV-T Pairs=Text-Only2026.02 | 89.6 | — | — | — | — | |
| VidVec-OModel Size=7B, Embedding=Optimized2026.02 | 89.6 | 98.5 | 99.3 | — | — | |
| InternVideo2s2-6BFinetuning=true2024.03 | 89.3 | — | — | — | — | |
| InternVideoBackbone=ViT-L/142022.12 | 87.2 | — | — | — | — | |
| InternVideoBackbone=ViT-L, inference protocol=dual-softmax2024.03 | 87.2 | — | — | — | — | |
| InternVideo-LModality=T-V, Finetuning=true, Evaluation frames=122024.12 | 87.2 | — | — | — | — | |
| InternVideo-Lsetting=finetuning, evaluation_frames=122025.07 | 87.2 | — | — | — | — | |
| UMT-LFinetuning=true2024.03 | 86 | — | — | — | — | |
| UMT-LModality=T-V, Finetuning and evaluation frames=122024.12 | 86 | — | — | — | — | |
| UMT-LModality=T-V, Finetuning=true, Evaluation frames=122024.12 | 86 | — | — | — | — | |
| UMT-Lsetting=finetuning, evaluation_frames=122025.07 | 86 | — | — | — | — | |
| VidLABackbone=CLIP-ViT-B/16, inference protocol=dual-softmax2024.03 | 85.6 | 99.8 | 100 | 95.1 | — | |
| PE-GV-Enc Params=1.9B, Video Data=22M, Resolution=448, Frame/ fps=8, Zero-shot Protocol=true2025.12 | 85.5 | — | — | — | — | |
| InternVideo2V-Enc Params=1.0B, Video Data=102M, Resolution=224, Frame/ fps=8, Zero-shot Protocol=true2025.12 | 85.4 | — | — | — | — | |
| InternVideo2-1BSetting=ZS2026.05 | 85.4 | — | — | — | — | |
| InternVideo2-6BV-T Pairs=100M2026.02 | 85.3 | — | — | — | — | |
| VidLABackbone=CLIP-ViT-B/32, inference protocol=dual-softmax2024.03 | 85 | 99.8 | 99.9 | 94.9 | — | |
| GRAM ModelModality=T-VA2024.12 | 84.6 | — | 100 | — | — | |
| GRAM ModelModality=T-VA, Finetuning=true2024.12 | 84.6 | — | — | — | — | |
| GRAM ModelModality=T-VAS2024.12 | 84.2 | — | 99.8 | — | — | |
| GRAM ModelModality=T-VAS, Finetuning=true2024.12 | 84.2 | — | — | — | — | |
| VASTModality=T-VA2024.12 | 84.1 | — | 99.8 | — | — | |
| VASTModality=T-VA, Finetuning=true2024.12 | 84.1 | — | — | — | — | |
| VASTModality=T-VAS2024.12 | 84 | — | 99.7 | — | — | |
| VASTModality=T-VAS, Finetuning=true2024.12 | 84 | — | — | — | — | |
| PMRLsetting=finetuning2025.07 | 83.4 | — | — | — | — | |
| B3Model Size=7B2026.02 | 83.1 | 96.4 | 98.5 | — | — | |
| LanguageBindV-Enc Params=0.3B, Video Data=10M, Resolution=224, Frame/ fps=16, Zero-shot Protocol=true2025.12 | 83.1 | — | — | — | — | |
| GRAMModality=T-VAS, Zero-shot=true2024.12 | 82.7 | — | — | — | — | |
| GRAM ModelModality=T-VAS, Evaluation Protocol=Zero-shot2024.12 | 82.7 | — | 98.1 | — | — | |
| LamRAModel Size=7B2026.02 | 82.5 | 95.4 | 97.7 | — | — | |
| GRAM ModelModality=T-V2024.12 | 81.6 | — | 99.8 | — | — | |
| GRAM ModelModality=T-V, Finetuning=true2024.12 | 81.6 | — | — | — | — | |
| UniME-V2Model Size=7B2026.02 | 81.5 | 96.1 | 98.4 | — | — | |
| VASTsetting=finetuning2025.07 | 81.3 | — | — | — | — | |
| VidLABackbone=CLIP-ViT-B/162024.03 | 81.1 | 99.5 | 99.9 | 93.5 | — | |
| MMRet-v1.5Model Size=7B2026.02 | 81.1 | 95.7 | 97.7 | — | — | |
| GRAMsetting=finetuning2025.07 | 80.6 | — | — | — | — | |
| PE-LV-Enc Params=0.3B, Video Data=22M, Resolution=336, Frame/ fps=8, Zero-shot Protocol=true2025.12 | 80.5 | — | — | — | — | |
| VLM2VecModel Size=7B2026.02 | 80.2 | 95.5 | 98.7 | — | — | |
| DRLBackbone=CLIP-ViT-B/162024.03 | 80.1 | 98.5 | 99.5 | 92.7 | — | |
| VLM2Vec-V2Model Size=7B2026.02 | 79.9 | 95.9 | 98.2 | — | — | |
| VidLABackbone=CLIP-ViT-B/322024.03 | 79.8 | 99.5 | 99.8 | 93 | — | |
| GRAMModality=T-VA, Zero-shot=true2024.12 | 79.2 | — | — | — | — | |
| GRAM ModelModality=T-VA, Evaluation Protocol=Zero-shot2024.12 | 79.2 | — | 99 | — | — | |
| GRAMModality=T-V, Zero-shot=true2024.12 | 79 | — | — | — | — | |
| GRAM ModelModality=T-V, Evaluation Protocol=Zero-shot2024.12 | 79 | — | 98.3 | — | — | |
| UNITEModel Size=7B2026.02 | 78.9 | 95.3 | 98.1 | — | — | |
| VASTModality=T-VAS, Zero-shot=true2024.12 | 78.7 | — | — | — | — | |
| VASTModality=T-VAS, Evaluation Protocol=Zero-shot2024.12 | 78.7 | — | 97.7 | — | — | |
| CLIP4Clip2022.12 | 78.3 | — | — | — | — | |
| CLIP4ClipModality=T-V2024.12 | 78.3 | — | — | — | — | |
| CLIP4ClipModality=T-V, Finetuning=true2024.12 | 78.3 | — | — | — | — | |
| CLIP4Clipsetting=finetuning2025.07 | 78.3 | — | — | — | — | |
| VASTModality=T-VA, Zero-shot=true2024.12 | 77.3 | — | — | — | — | |
| VideoPrism-bModality=T-V, Zero-shot=true2024.12 | 77.1 | — | — | — | — | |
| VASTModality=T-VA, Evaluation Protocol=Zero-shot2024.12 | 77.1 | — | 95.2 | — | — | |
| VideoPrism-gV-T Pairs=600M2026.02 | 77.1 | — | — | — | — | |
| VideoPrism-bZero-shot=true2025.07 | 77.1 | — | — | — | — | |
| DRLBackbone=CLIP-ViT-B/322024.03 | 77 | 98 | 99.4 | 91.5 | — | |
| VideoPrism-bModality=T-V, Evaluation Protocol=Zero-shot2024.12 | 76.2 | — | — | — | — | |
| PMRLZero-shot=true2025.07 | 75.2 | — | — | — | — | |
| VASTZero-shot=true2025.07 | 74.8 | — | — | — | — | |
| GRAMZero-shot=true2025.07 | 74.7 | — | — | — | — | |
| VideoCoCaModality=T-V, Zero-shot=true2024.12 | 73.6 | — | — | — | — | |
| VideoCoCaZero-shot=true2025.07 | 73.6 | — | — | — | — | |
| CLIP4ClipSetting=FT2026.05 | 73.4 | — | — | — | — | |
| ImageBindV-Enc Params=0.6B, Video Data=3M, Resolution=224, Frame/ fps=16, Zero-shot Protocol=true2025.12 | 69.8 | — | — | — | — | |
| InternVideoZero-shot=true2022.12 | 69.5 | — | — | — | — | |
| InternVideo-LModality=T-V, Zero-shot=true2024.12 | 69.5 | — | — | — | — | |
| InternVideo-LModality=T-V, Evaluation Protocol=Zero-shot2024.12 | 69.5 | — | — | — | — | |
| InternVideo-LV-T Pairs=12M2026.02 | 69.5 | — | — | — | — | |
| InternVideo-LZero-shot=true2025.07 | 69.5 | — | — | — | — | |
| SigLIP2-g-optV-Enc Params=1.1B, Resolution=384, Frame/ fps=8, Zero-shot Protocol=true2025.12 | 68.2 | — | — | — | — | |
| Gemini Embedding 2Setting=ZS2026.05 | 66.73 | — | — | — | — | |
| SigLIP2-L/16V-Enc Params=0.3B, Resolution=384, Frame/ fps=8, Zero-shot Protocol=true2025.12 | 64.3 | — | — | — | — | |
| CLIPZero-shot=true2022.12 | 59.2 | — | — | — | — | |
| OmniRetriever-7BSetting=ZS2026.05 | 54.97 | — | — | — | — | |
| AutoMergeEvaluation Protocol=Zero-shot, Training Samples=4M2025.05 | — | 58.1 | — | — | — | |
| GME 7BModel Size=7B2025.09 | — | — | — | — | 33 | |
| InternVideo-LZero-shot=true2024.03 | — | — | — | 69.5 | — | |
| InternVideo2_s2-1BZero-shot=true2024.03 | — | — | — | 85.4 | — | |
| InternVideo2_s2-6BZero-shot=true2024.03 | — | — | — | 85.3 | — | |
| RLTEvaluation Protocol=Zero-shot, Training Samples=4M2025.05 | — | 58.4 | — | — | — | |
| RLTmode=Zero-shot2026.02 | — | 58.4 | — | — | — | |
| TokenLearnerEvaluation Protocol=Zero-shot, Training Samples=4M2025.05 | — | 58.5 | — | — | — | |
| TokenLearnermode=Zero-shot2026.02 | — | 58.8 | — | — | — | |
| ToMeEvaluation Protocol=Zero-shot, Training Samples=4M2025.05 | — | 56.6 | — | — | — | |
| TrajViTEvaluation Protocol=Zero-shot, Training Samples=4M2025.05 | — | 61 | — | — | — | |
| TrajViTmode=Zero-shot2026.02 | — | 61.1 | — | — | — | |
| TrajViT-2mode=Zero-shot2026.02 | — | 65 | — | — | — | |
| VideoCoCa-gZero-shot=true2024.03 | — | — | — | 73.6 | — | |
| VideoPrism-gZero-shot=true2024.03 | — | — | — | 77.1 | — | |
| ViT3DEvaluation Protocol=Zero-shot, Training Samples=4M2025.05 | — | 59.1 | — | — | — |