Video Question Answering on MSRVTT-QA (test)
88.2AccuracyCLIPBERT
Evaluation Results
| Method | Links | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| CLIPBERTTraining input sampling (Ntrain x T)=8x22021.02 | 88.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CLIPBERTTraining input sampling (Ntrain x T)=4x12021.02 | 87.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ActBERTPre-training=PT (HowTo100M)2021.02 | 85.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| JSFusion2021.02 | 83.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MLB2021.02 | 76.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| PLLaVALLM Size=34B, Vision Encoder=CLIP-L, Training Mode=Fine-tuned2024.07 | 68.7 | — | — | — | — | — | — | — | — | — | 3.8 | — | — | — | |
| SF-LLaVA-34BLLM Size=34B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 67.4 | — | — | — | — | — | — | — | — | — | 3.7 | — | — | — | |
| CT-SAN2021.02 | 66.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ST-VQA2021.02 | 66.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SF-LLaVA-7BLLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 65.8 | — | — | — | — | — | — | — | — | — | 3.6 | — | — | — | |
| SNUVL2021.02 | 65.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| IG-VLM (GPT-4V)Vision Encoder=Unknown, LLM Size=GPT-4, Inference Vision=single, Inference LLM=single, Video Trained=X2024.03 | 63.8 | — | — | — | — | — | — | — | — | 3.5 | — | — | — | — | |
| IG-VLM (LLaVA v1.6)Vision Encoder=ViT-L, LLM Size=7B, Inference Vision=single, Inference LLM=single, Video Trained=X2024.03 | 63.7 | — | — | — | — | — | — | — | — | 3.5 | — | — | — | — | |
| IG-VLMLLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 63.7 | — | — | — | — | — | — | — | — | — | 3.5 | — | — | — | |
| ST-LLMVision Size=1.3B, Modality=V, Pretrain Data=null, Finetune Data=400K2024.12 | 63.2 | — | — | — | — | — | — | — | — | — | 3.4 | — | — | — | |
| IG-VLM (CogAgent)Vision Encoder=CLIP-E, LLM Size=7B, Inference Vision=single, Inference LLM=single, Video Trained=X2024.03 | 62.7 | — | — | — | — | — | — | — | — | 3.6 | — | — | — | — | |
| IG-VLM (LLaVA v1.6)Vision Encoder=ViT-L, LLM Size=13B, Inference Vision=single, Inference LLM=single, Video Trained=X2024.03 | 62.6 | — | — | — | — | — | — | — | — | 3.4 | — | — | — | — | |
| IG-VLM (LLaVA v1.6)Vision Encoder=ViT-L, LLM Size=34B, Inference Vision=single, Inference LLM=single, Video Trained=X2024.03 | 62.4 | — | — | — | — | — | — | — | — | 3.5 | — | — | — | — | |
| IG-VLMLLM Size=34B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 62.4 | — | — | — | — | — | — | — | — | — | 3.5 | — | — | — | |
| PLLaVALLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Fine-tuned2024.07 | 62 | — | — | — | — | — | — | — | — | — | 3.5 | — | — | — | |
| VideoGPT+LLM Size=3.8B, Vision Encoder=CLIP-L, Training Mode=Fine-tuned2024.07 | 60.6 | — | — | — | — | — | — | — | — | — | 3.6 | — | — | — | |
| Vista-LLaMAVision Encoder=CLIP-G, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=O2024.03 | 60.5 | — | — | — | — | — | — | — | — | 3.3 | — | — | — | — | |
| Vista-LLaMALLM Size=7B, Vision Encoder=CLIP-G, Training Mode=Fine-tuned2024.07 | 60.5 | — | — | — | — | — | — | — | — | — | 3.3 | — | — | — | |
| MiCo-Chat-7BLLM=Vicuna-7B, Res.=224, Zero-shot=true2024.06 | 60.1 | — | — | — | — | — | — | — | — | — | 3.6 | — | — | — | |
| FreeVALLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 60 | — | — | — | — | — | — | — | — | — | 3.5 | — | — | — | |
| Video-LLaVAVision Encoder=ViT-L, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=O2024.03 | 59.2 | — | — | — | — | — | — | — | — | 3.5 | — | — | — | — | |
| Video-LLaVALLM Size=7B, Vision Encoder=ViT-L, Training Mode=Fine-tuned2024.07 | 59.2 | — | — | — | — | — | — | — | — | — | 3.5 | — | — | — | |
| Video-LLaVALLM size=7B2023.11 | 59.2 | — | — | — | — | — | — | — | — | — | 3.5 | — | — | — | |
| Video-LLaVAVision Size=425M, Modality=V+I, Pretrain Data=1.26M, Finetune Data=765K2024.12 | 59.2 | — | — | — | — | — | — | — | — | — | 3.5 | — | — | — | |
| LLAMA-VIDLLM=Vicuna-13B, Res.=224, Setting=zero-shot, Tokens per frame=22023.11 | 58.9 | — | — | — | — | — | — | — | — | 3.3 | — | — | — | — | |
| LLaMA-VIDVision Encoder=CLIP-G, LLM Size=13B, Inference Vision=multiple, Inference LLM=single, Video Trained=O2024.03 | 58.9 | — | — | — | — | — | — | — | — | 3.3 | — | — | — | — | |
| Video-LLaVAVision Size=425M, Modality=V, Pretrain Data=702K, Finetune Data=100K2024.12 | 58.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLAMA-VIDLLM=Vicuna-7B, Res.=224, Setting=zero-shot, Tokens per frame=22023.11 | 57.7 | — | — | — | — | — | — | — | — | 3.2 | — | — | — | — | |
| LLaMA-VIDLLM=Vicuna-7B, Res.=224, Zero-shot=true2024.06 | 57.7 | — | — | — | — | — | — | — | — | — | 3.2 | — | — | — | |
| LLaMA-VIDLLM Size=13B, Vision Encoder=CLIP-G, Training Mode=Fine-tuned2024.07 | 57.7 | — | — | — | — | — | — | — | — | — | 3.2 | — | — | — | |
| LLaMA-VIDVision Size=1B, Modality=V+I, Pretrain Data=790K, Finetune Data=763K2024.12 | 57.7 | — | — | — | — | — | — | — | — | — | 3.2 | — | — | — | |
| TGB (Vicuna7B)LLM size=7B, Zero-shot=true, Base Model=Vicuna7B2024.02 | 57.3 | — | — | — | — | — | — | — | — | — | 3.3 | — | — | — | |
| LLaVolta#Stages=Three, Scheme=compr. deeper, #Tokens=84776, CR=174%, TFLOPs=17.29, Train-time=37.1h2024.06 | 57.2 | — | — | — | — | — | — | — | — | — | 3.51 | — | — | — | |
| LLaVolta#Stages=Three, Scheme=compr. wider, #Tokens=83256, CR=177%, TFLOPs=16.86, Train-time=37.0h2024.06 | 57.2 | — | — | — | — | — | — | — | — | — | 3.51 | — | — | — | |
| LLaVolta#Stages=Four, Scheme=wider then deeper, #Tokens=88704, CR=166%, TFLOPs=18.32, Train-time=37.2h2024.06 | 57.2 | — | — | — | — | — | — | — | — | — | 3.51 | — | — | — | |
| BT-AdapterLLM=Vicuna-7B, Res.=-, Setting=zero-shot, Tokens per frame=22023.11 | 57 | — | — | — | — | — | — | — | — | 3.2 | — | — | — | — | |
| BT-Adapterinstruction tuning=true2023.09 | 57 | — | — | — | — | — | — | — | — | — | 3.2 | — | — | — | |
| LLaVolta#Stages=Two, Scheme=compression, #Tokens=80496, CR=183%, TFLOPs=17.73, Train-time=37.1h2024.06 | 56.9 | — | — | — | — | — | — | — | — | — | 3.5 | — | — | — | |
| LLaVolta#Stages=Four, Scheme=deeper then wider, #Tokens=86904, CR=170%, TFLOPs=18.64, Train-time=37.1h2024.06 | 56.9 | — | — | — | — | — | — | — | — | — | 3.49 | — | — | — | |
| VideoLLaVA#Stages=Single, Scheme=no compression, #Tokens=147456, TFLOPs=29.88, Train-time=40.7h2024.06 | 56.8 | — | — | — | — | — | — | — | — | — | 3.48 | — | — | — | |
| Full KV CachingModel=VideoLLaVA-7B, Cache Memory (avg. per instance) Size (MB)=1114.6, Cache Memory Reduction=1×, Batch Inf. Average Change=−2026.03 | 55.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| FastVModel=VideoLLaVA-7B, Params=k = 5, e = 50%, Cache Memory (avg. per instance) Size (MB)=1114.6, Cache Memory Reduction=1×, Batch Inf. Average Change=+23%2026.03 | 55.52 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| AttentionPackModel=VideoLLaVA-7B, Params=Rkv = Rvv = 128, Cache Memory (avg. per instance) Size (MB)=137.5, Cache Memory Reduction=8.11×, Batch Inf. Average Change=+60%2026.03 | 55.47 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Chat-UniViLLM Size=7B, Zero-shot=true2023.11 | 55 | — | — | — | — | — | — | — | — | 3.1 | — | — | — | — | |
| Video-PandaVision Size=45M, Modality=V, Pretrain Data=702K, Finetune Data=100K2024.12 | 54.8 | — | — | — | — | — | — | — | — | — | 3.4 | — | — | — | |
| H2OModel=VideoLLaVA-7B, Params=e = 50%, Cache Memory (avg. per instance) Size (MB)=557.3, Cache Memory Reduction=2×, Batch Inf. Average Change=+32%2026.03 | 54.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Chat-UniViVision Encoder=ViT-L, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=O2024.03 | 54.6 | — | — | — | — | — | — | — | — | 3.1 | — | — | — | — | |
| Chat-UniViLLM size=7B2023.11 | 54.6 | — | — | — | — | — | — | — | — | — | 3.1 | — | — | — | |
| ChatUniViVision Size=307M, Modality=V+I, Pretrain Data=1.6M, Finetune Data=649K2024.12 | 54.6 | — | — | — | — | — | — | — | — | — | 3.1 | — | — | — | |
| ScissorhandsModel=VideoLLaVA-7B, Params=e = 50%, Cache Memory (avg. per instance) Size (MB)=557.3, Cache Memory Reduction=2×, Batch Inf. Average Change=+32%2026.03 | 54.33 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VideoChat2Vision Encoder=UMT-L, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=O2024.03 | 54.1 | — | — | — | — | — | — | — | — | 3.3 | — | — | — | — | |
| VideoChat2LLM=Vicuna-7B, Res.=224, Zero-shot=true2024.06 | 54.1 | — | — | — | — | — | — | — | — | — | 3.3 | — | — | — | |
| VideoChat2LLM Size=7B, Vision Encoder=UMT-L, Training Mode=Fine-tuned2024.07 | 54.1 | — | — | — | — | — | — | — | — | — | 3.3 | — | — | — | |
| VideoChat2Vision Size=496M, Modality=V+I, Pretrain Data=37M, Finetune Data=2M2024.12 | 54.1 | — | — | — | — | — | — | — | — | — | 3.3 | — | — | — | |
| MovieChat+Evaluator=GPT-3.52024.04 | 53.9 | — | — | — | — | — | — | — | — | — | 2.7 | — | — | — | |
| MovieChat+LLM Size=7B, Vision Encoder=CLIP-G, Training Mode=Fine-tuned2024.07 | 53.9 | — | — | — | — | — | — | — | — | — | 2.7 | — | — | — | |
| TGB (BLIP2)LLM size=3B, Zero-shot=true, Base Model=BLIP22024.02 | 53.5 | — | — | — | — | — | — | — | — | — | 3.1 | — | — | — | |
| MovieChat2023.07 | 52.7 | — | — | — | — | — | — | — | — | 2.6 | — | — | — | — | |
| MovieChatVision Encoder=CLIP-G, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=O2024.03 | 52.7 | — | — | — | — | — | — | — | — | 2.6 | — | — | — | — | |
| MovieChatEvaluator=GPT-3.52024.04 | 52.7 | — | — | — | — | — | — | — | — | — | 2.6 | — | — | — | |
| MovieChatLLM Size=7B, Vision Encoder=CLIP-G, Training Mode=Fine-tuned2024.07 | 52.7 | — | — | — | — | — | — | — | — | — | 2.6 | — | — | — | |
| BT-Adapterinstruction tuning=false2023.09 | 51.2 | — | — | — | — | — | — | — | — | — | 2.9 | — | — | — | |
| Mirasol3BEvaluation Protocol=Open-ended generation2023.11 | 50.42 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Mirasol3B - TTMEvaluation Protocol=Open-ended generation, Combiner=TTM2023.11 | 50.01 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| FlamingoShots=32, Evaluation Protocol=In-context learning2022.04 | 49.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| EVE*Vision Size=30M, Modality=V, Pretrain Data=702K, Finetune Data=100K2024.12 | 49.7 | — | — | — | — | — | — | — | — | — | 3 | — | — | — | |
| MaMMUTsetting=open-ended generation, pre-training=image-text pre-training only2023.03 | 49.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Video-ChatGPTLLM=Vicuna-7B, Res.=224, Setting=zero-shot, Tokens per frame=22023.11 | 49.3 | — | — | — | — | — | — | — | — | 2.8 | — | — | — | — | |
| Video-ChatGPT2023.07 | 49.3 | — | — | — | — | — | — | — | — | 2.8 | — | — | — | — | |
| Video-ChatGPTVision Encoder=ViT-L, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=O2024.03 | 49.3 | — | — | — | — | — | — | — | — | 2.8 | — | — | — | — | |
| Video-ChatGPTLLM Size=7B, Zero-shot=true2023.11 | 49.3 | — | — | — | — | — | — | — | — | 2.8 | — | — | — | — | |
| Video-ChatGPTEvaluator=GPT-3.52024.04 | 49.3 | — | — | — | — | — | — | — | — | — | 2.8 | — | — | — | |
| Video-ChatGPTZero-shot=true, Parameters=7B2023.06 | 49.3 | — | — | — | — | — | — | — | — | — | 2.8 | — | — | — | |
| Video-ChatGPTLLM=Vicuna-7B, Res.=224, Zero-shot=true2024.06 | 49.3 | — | — | — | — | — | — | — | — | — | 2.8 | — | — | — | |
| VideoChatGPTinstruction tuning=false2023.09 | 49.3 | — | — | — | — | — | — | — | — | — | 2.8 | — | — | — | |
| Video-ChatGPTLLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Fine-tuned2024.07 | 49.3 | — | — | — | — | — | — | — | — | — | 2.8 | — | — | — | |
| Video-ChatGPTLLM size=7B2023.11 | 49.3 | — | — | — | — | — | — | — | — | — | 2.8 | — | — | — | |
| Video-ChatGPTLLM size=7B, Zero-shot=true2024.02 | 49.3 | — | — | — | — | — | — | — | — | — | 2.8 | — | — | — | |
| Video-ChatGPTVision Size=307M, Modality=V, Pretrain Data=100K, Finetune Data=100K2024.12 | 49.3 | — | — | — | — | — | — | — | — | — | 2.8 | — | — | — | |
| mPLUG-2Pre-training Data Size=17M2023.02 | 48 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| M-PLUG2Evaluation Protocol=Classification2023.11 | 48 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MuLTI-L#PT=5.5M2023.03 | 47.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Flamingomode=Fine-tuned2022.04 | 47.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| FlamingoEvaluation Protocol=Fine-tuned2022.04 | 47.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| FlamingoPre-training Data Size=2.3B2023.02 | 47.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Flamingosetting=open-ended generation, pre-training=image-text2023.03 | 47.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Flamingo#PT=2139M2023.03 | 47.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| FlamingoEvaluation Protocol=Classification2023.11 | 47.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Flamingo2022.05 | 47.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVideoPre-training Data Size=12M2023.02 | 47.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVideosetting=open-ended generation, pre-training=image-text2023.03 | 47.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVideo#Pairs=646M, GFLOPS=666.22024.03 | 47.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| UMT-L#Pairs=25M, GFLOPS=984.62024.03 | 47.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| UMT-LEvaluation Protocol=Classification2023.11 | 47.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVideoEvaluation Protocol=Classification2023.11 | 47.1 | — | — | — | — | — | — | — | — | — | — | — | — | — |