Video Question Answering on MSVD-QA (test)
87.8AccuracyVideo-QTR
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Video-QTRModel Category=Our Method2025.12 | 87.8 | — | — | — | — | — | — | 4.63 | — | — | |
| gpt-4.1Model Category=Large Multimodal Models2025.12 | 85.15 | — | — | — | — | — | — | 4.56 | — | — | |
| qwen2.5-vl-maxModel Category=Large Multimodal Models2025.12 | 84.14 | — | — | — | — | — | — | 4.51 | — | — | |
| Flash-VstreamModel Category=Video QA Methods2025.12 | 80.3 | — | — | — | — | — | — | 3.9 | — | — | |
| gemini2.5-proModel Category=Large Multimodal Models2025.12 | 80.2 | — | — | — | — | — | — | 3.9 | — | — | |
| PLLaVALLM Size=34B, Vision Encoder=CLIP-L, Training Mode=Fine-tuned2024.07 | 79.9 | — | — | — | — | — | — | 4.2 | — | — | |
| SF-LLaVA-34BLLM Size=34B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 79.9 | — | — | — | — | — | — | 4.1 | — | — | |
| IG-VLM (LLaVA v1.6)Vision Encoder=ViT-L, LLM Size=34B, Inference Vision=single, Inference LLM=single, Video Trained=X2024.03 | 79.6 | — | — | — | — | — | — | 4.1 | — | — | |
| IG-VLMLLM Size=34B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 79.6 | — | — | — | — | — | — | 4.1 | — | — | |
| SF-LLaVA-7BLLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 79.1 | — | — | — | — | — | — | 4.1 | — | — | |
| IG-VLM (LLaVA v1.6)Vision Encoder=ViT-L, LLM Size=7B, Inference Vision=single, Inference LLM=single, Video Trained=X2024.03 | 78.8 | — | — | — | — | — | — | 4.1 | — | — | |
| IG-VLMLLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 78.8 | — | — | — | — | — | — | 4.1 | — | — | |
| IG-VLM (LLaVA v1.6)Vision Encoder=ViT-L, LLM Size=13B, Inference Vision=single, Inference LLM=single, Video Trained=X2024.03 | 77.4 | — | — | — | — | — | — | 4.1 | — | — | |
| IG-VLM (CogAgent)Vision Encoder=CLIP-E, LLM Size=7B, Inference Vision=single, Inference LLM=single, Video Trained=X2024.03 | 76.7 | — | — | — | — | — | — | 4.1 | — | — | |
| IG-VLMModel Category=Video QA Methods2025.12 | 76.7 | — | — | — | — | — | — | 4.1 | — | — | |
| PLLaVALLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Fine-tuned2024.07 | 76.6 | — | — | — | — | — | — | 4.1 | — | — | |
| PLLaVAModel Category=Video QA Methods2025.12 | 76.6 | — | — | — | — | — | — | 4.1 | — | — | |
| MovieChat+Evaluator=GPT-3.52024.04 | 76.5 | — | — | — | — | — | — | 3.9 | — | — | |
| MovieChat+LLM Size=7B, Vision Encoder=CLIP-G, Training Mode=Fine-tuned2024.07 | 76.5 | — | — | — | — | — | — | 3.9 | — | — | |
| IG-VLM (GPT-4V)Vision Encoder=Unknown, LLM Size=GPT-4, Inference Vision=single, Inference LLM=single, Video Trained=X2024.03 | 76.3 | — | — | — | — | — | — | 4 | — | — | |
| DeepStack-LLLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 76 | — | — | — | — | — | — | 4 | — | — | |
| EMA2026.02 | 75.8 | — | — | — | — | — | — | 4.1 | — | — | |
| MovieChat2023.07 | 75.2 | — | — | — | — | — | — | 3.8 | — | — | |
| MovieChatVision Encoder=CLIP-G, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=O2024.03 | 75.2 | — | — | — | — | — | — | 3.8 | — | — | |
| MovieChatEvaluator=GPT-3.52024.04 | 75.2 | — | — | — | — | — | — | 3.8 | — | — | |
| MovieChatLLM Size=7B, Vision Encoder=CLIP-G, Training Mode=Fine-tuned2024.07 | 75.2 | — | — | — | — | — | — | 3.8 | — | — | |
| MovieChatModel Category=Video QA Methods2025.12 | 75.2 | — | — | — | — | — | — | 3.8 | — | — | |
| MovieChat2026.02 | 75.2 | — | — | — | — | — | — | 3.8 | — | — | |
| ST-LLMVision Size=1.3B, Modality=V, Pretrain Data=null, Finetune Data=400K2024.12 | 74.6 | — | — | — | — | — | — | 3.9 | — | — | |
| FreeVALLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 73.8 | — | — | — | — | — | — | 4.1 | — | — | |
| MiCo-Chat-7BLLM=Vicuna-7B, Res.=224, Zero-shot=true2024.06 | 73.7 | — | — | — | — | — | — | 4.1 | — | — | |
| Video-LaVIT2026.02 | 73.2 | — | — | — | — | — | — | 3.9 | — | — | |
| ReMoRa2026.02 | 73.1 | — | — | — | — | — | — | 4 | — | — | |
| VideoGPT+LLM Size=3.8B, Vision Encoder=CLIP-L, Training Mode=Fine-tuned2024.07 | 72.4 | — | — | — | — | — | — | 3.9 | — | — | |
| TGB (Vicuna7B)LLM size=7B, Zero-shot=true, Base Model=Vicuna7B2024.02 | 71.4 | — | — | — | — | — | — | 3.9 | — | — | |
| Video-LLaMA2LLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Fine-tuned2024.07 | 70.9 | — | — | — | — | — | — | 3.8 | — | — | |
| Video-LLaVAVision Encoder=ViT-L, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=O2024.03 | 70.7 | — | — | — | — | — | — | 3.9 | — | — | |
| Video-LLaVALLM Size=7B, Vision Encoder=ViT-L, Training Mode=Fine-tuned2024.07 | 70.7 | — | — | — | — | — | — | 3.9 | — | — | |
| Video-LLaVALLM size=7B2023.11 | 70.7 | — | — | — | — | — | — | 3.9 | — | — | |
| Video-LLaVAVision Size=425M, Modality=V+I, Pretrain Data=1.26M, Finetune Data=765K2024.12 | 70.7 | — | — | — | — | — | — | 3.9 | — | — | |
| Video-LLaVA2026.02 | 70.7 | — | — | — | — | — | — | 3.9 | — | — | |
| Video-LLaMA2LLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Fine-tuned2024.07 | 70.5 | — | — | — | — | — | — | 3.8 | — | — | |
| LLAMA-VIDLLM=Vicuna-13B, Res.=224, Setting=zero-shot, Tokens per frame=22023.11 | 70 | — | — | — | — | — | — | 3.7 | — | — | |
| VideoChat2Vision Encoder=UMT-L, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=O2024.03 | 70 | — | — | — | — | — | — | 3.9 | — | — | |
| LLaMA-VIDVision Encoder=CLIP-G, LLM Size=13B, Inference Vision=multiple, Inference LLM=single, Video Trained=O2024.03 | 70 | — | — | — | — | — | — | 3.7 | — | — | |
| VideoChat2LLM=Vicuna-7B, Res.=224, Zero-shot=true2024.06 | 70 | — | — | — | — | — | — | 3.9 | — | — | |
| VideoChat2LLM Size=7B, Vision Encoder=UMT-L, Training Mode=Fine-tuned2024.07 | 70 | — | — | — | — | — | — | 3.9 | — | — | |
| VideoChat2Vision Size=496M, Modality=V+I, Pretrain Data=37M, Finetune Data=2M2024.12 | 70 | — | — | — | — | — | — | 3.9 | — | — | |
| LLaVolta#Stages=Four, Scheme=deeper then wider, #Tokens=86904, CR=170%, TFLOPs=18.64, Train-time=37.1h2024.06 | 69.8 | — | — | — | — | — | — | 3.74 | — | — | |
| LLAMA-VIDLLM=Vicuna-7B, Res.=224, Setting=zero-shot, Tokens per frame=22023.11 | 69.7 | — | — | — | — | — | — | 3.7 | — | — | |
| LLaMA-VIDLLM=Vicuna-7B, Res.=224, Zero-shot=true2024.06 | 69.7 | — | — | — | — | — | — | 3.7 | — | — | |
| LLaMA-VIDLLM Size=13B, Vision Encoder=CLIP-G, Training Mode=Fine-tuned2024.07 | 69.7 | — | — | — | — | — | — | 3.7 | — | — | |
| LLaMA-VIDVision Size=1B, Modality=V+I, Pretrain Data=790K, Finetune Data=763K2024.12 | 69.7 | — | — | — | — | — | — | 3.7 | — | — | |
| LLaMA-VID2026.02 | 69.7 | — | — | — | — | — | — | 3.7 | — | — | |
| FastVModel=VideoLLaVA-7B, Params=k = 5, e = 50%, Cache Memory (avg. per instance) Size (MB)=1114.6, Cache Memory Reduction=1×, Batch Inf. Average Change=+23%2026.03 | 69.6 | — | — | — | — | — | — | — | — | — | |
| Full KV CachingModel=VideoLLaVA-7B, Cache Memory (avg. per instance) Size (MB)=1114.6, Cache Memory Reduction=1×, Batch Inf. Average Change=−2026.03 | 69.33 | — | — | — | — | — | — | — | — | — | |
| Chat-UniViLLM Size=7B, Zero-shot=true2023.11 | 69.3 | — | — | — | — | — | — | 3.7 | — | — | |
| LLaVolta#Stages=Three, Scheme=compr. deeper, #Tokens=84776, CR=174%, TFLOPs=17.29, Train-time=37.1h2024.06 | 69.3 | — | — | — | — | — | — | 3.73 | — | — | |
| AttentionPackModel=VideoLLaVA-7B, Params=Rkv = Rvv = 128, Cache Memory (avg. per instance) Size (MB)=137.5, Cache Memory Reduction=8.11×, Batch Inf. Average Change=+60%2026.03 | 69.21 | — | — | — | — | — | — | — | — | — | |
| VideoLLaVA#Stages=Single, Scheme=no compression, #Tokens=147456, TFLOPs=29.88, Train-time=40.7h2024.06 | 69.1 | — | — | — | — | — | — | 3.69 | — | — | |
| LLaVolta#Stages=Four, Scheme=wider then deeper, #Tokens=88704, CR=166%, TFLOPs=18.32, Train-time=37.2h2024.06 | 69.1 | — | — | — | — | — | — | 3.72 | — | — | |
| LLaVolta#Stages=Two, Scheme=compression, #Tokens=80496, CR=183%, TFLOPs=17.73, Train-time=37.1h2024.06 | 69 | — | — | — | — | — | — | 3.71 | — | — | |
| LLaVolta#Stages=Three, Scheme=compr. wider, #Tokens=83256, CR=177%, TFLOPs=16.86, Train-time=37.0h2024.06 | 69 | — | — | — | — | — | — | 3.72 | — | — | |
| BT-AdapterLLM=Vicuna-7B, Res.=-, Setting=zero-shot, Tokens per frame=22023.11 | 67.5 | — | — | — | — | — | — | 3.7 | — | — | |
| BT-Adapterinstruction tuning=true2023.09 | 67.5 | — | — | — | — | — | — | 3.7 | — | — | |
| BT-Adapter2026.02 | 67.5 | — | — | — | — | — | — | 3.7 | — | — | |
| BT-Adapterinstruction tuning=false2023.09 | 67 | — | — | — | — | — | — | 3.6 | — | — | |
| H2OModel=VideoLLaVA-7B, Params=e = 50%, Cache Memory (avg. per instance) Size (MB)=557.3, Cache Memory Reduction=2×, Batch Inf. Average Change=+32%2026.03 | 66.54 | — | — | — | — | — | — | — | — | — | |
| TGB (BLIP2)LLM size=3B, Zero-shot=true, Base Model=BLIP22024.02 | 66 | — | — | — | — | — | — | 3.6 | — | — | |
| ScissorhandsModel=VideoLLaVA-7B, Params=e = 50%, Cache Memory (avg. per instance) Size (MB)=557.3, Cache Memory Reduction=2×, Batch Inf. Average Change=+32%2026.03 | 65.9 | — | — | — | — | — | — | — | — | — | |
| Vista-LLaMAVision Encoder=CLIP-G, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=O2024.03 | 65.3 | — | — | — | — | — | — | 3.6 | — | — | |
| Vista-LLaMALLM Size=7B, Vision Encoder=CLIP-G, Training Mode=Fine-tuned2024.07 | 65.3 | — | — | — | — | — | — | 3.6 | — | — | |
| Chat-UniViVision Encoder=ViT-L, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=O2024.03 | 65 | — | — | — | — | — | — | 3.6 | — | — | |
| Chat-UniViLLM size=7B2023.11 | 65 | — | — | — | — | — | — | 3.6 | — | — | |
| ChatUniViVision Size=307M, Modality=V+I, Pretrain Data=1.6M, Finetune Data=649K2024.12 | 65 | — | — | — | — | — | — | 3.6 | — | — | |
| Chat-UniVi2026.02 | 65 | — | — | — | — | — | — | 3.6 | — | — | |
| Video-ChatGPTLLM=Vicuna-7B, Res.=224, Setting=zero-shot, Tokens per frame=22023.11 | 64.9 | — | — | — | — | — | — | 3.3 | — | — | |
| Video-ChatGPT2023.07 | 64.9 | — | — | — | — | — | — | 3.3 | — | — | |
| Video-ChatGPTVision Encoder=ViT-L, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=O2024.03 | 64.9 | — | — | — | — | — | — | 3.3 | — | — | |
| Video-ChatGPTLLM Size=7B, Zero-shot=true2023.11 | 64.9 | — | — | — | — | — | — | 3.3 | — | — | |
| Video-ChatGPTEvaluator=GPT-3.52024.04 | 64.9 | — | — | — | — | — | — | 3.3 | — | — | |
| Video-ChatGPTZero-shot=true, Parameters=7B2023.06 | 64.9 | — | — | — | — | — | — | 3.3 | — | — | |
| Video-ChatGPTLLM=Vicuna-7B, Res.=224, Zero-shot=true2024.06 | 64.9 | — | — | — | — | — | — | 3.3 | — | — | |
| VideoChatGPTinstruction tuning=false2023.09 | 64.9 | — | — | — | — | — | — | 3.3 | — | — | |
| Video-ChatGPTLLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Fine-tuned2024.07 | 64.9 | — | — | — | — | — | — | 3.3 | — | — | |
| Video-ChatGPTLLM size=7B2023.11 | 64.9 | — | — | — | — | — | — | 3.3 | — | — | |
| Video-ChatGPTLLM size=7B, Zero-shot=true2024.02 | 64.9 | — | — | — | — | — | — | 3.3 | — | — | |
| Video-ChatGPTVision Size=307M, Modality=V, Pretrain Data=100K, Finetune Data=100K2024.12 | 64.9 | — | — | — | — | — | — | 3.3 | — | — | |
| Video-ChatGPTModel Category=Video QA Methods2025.12 | 64.9 | — | — | — | — | — | — | 3.3 | — | — | |
| Video-ChatGPT2026.02 | 64.9 | — | — | — | — | — | — | 3.3 | — | — | |
| Video-ChatGPTZero-shot=true2024.05 | 64.9 | — | — | — | — | — | — | 3.3 | — | — | |
| Video-LLaVAVision Size=425M, Modality=V, Pretrain Data=702K, Finetune Data=100K2024.12 | 64.8 | — | — | — | — | — | — | — | — | — | |
| Video-PandaVision Size=45M, Modality=V, Pretrain Data=702K, Finetune Data=100K2024.12 | 64.7 | — | — | — | — | — | — | 3.8 | — | — | |
| EVE*Vision Size=30M, Modality=V, Pretrain Data=702K, Finetune Data=100K2024.12 | 60.5 | — | — | — | — | — | — | 3.3 | — | — | |
| MaMMUTsetting=open-ended generation, pre-training=image-text pre-training only2023.03 | 60.2 | — | — | — | — | — | — | — | — | — | |
| GIT2Pre-training Data Size=12.9B2023.02 | 58.2 | — | — | — | — | — | — | — | — | — | |
| GIT2setting=open-ended generation, pre-training=image-text2023.03 | 58.2 | — | — | — | — | — | — | — | — | — | |
| GIT22022.05 | 58.2 | — | — | — | — | — | — | — | — | — | |
| mPLUG-2Pre-training Data Size=17M2023.02 | 58.1 | — | — | — | — | — | — | — | — | — | |
| VideoCoCaPre-training Data Size=3B2023.02 | 56.9 | — | — | — | — | — | — | — | — | — |