Video Question Answering on ActivityNet-QA (test)
82.78Accuracygpt-4.1
Evaluation Results
| Method | Links | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| gpt-4.1Model Category=Large Multimodal Models2025.12 | 82.78 | — | — | — | — | — | — | — | 4.13 | — | — | — | — | |
| Video-QTRModel Category=Our Method2025.12 | 82.32 | — | — | — | — | — | — | — | 4.31 | — | — | — | — | |
| gemini2.5-proModel Category=Large Multimodal Models2025.12 | 79.7 | — | — | — | — | — | — | — | 3.88 | — | — | — | — | |
| qwen2.5-vl-maxModel Category=Large Multimodal Models2025.12 | 77.32 | — | — | — | — | — | — | — | 4.01 | — | — | — | — | |
| LLaVA-NeXT-Video-DPOLLM Size=34B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 64.4 | 3.6 | — | — | — | — | — | — | — | — | — | — | — | |
| Regular Sampling + LLaVA + LLMSampling Strategy=Regular Sampling, LMM=LLaVA-1.5, LLM=Vicuna-v1.5, Fusion=Late fusion2026.01 | 63.3 | — | — | — | — | — | — | — | 3.87 | — | — | — | — | |
| GPT-4oZero-shot=true2024.07 | 61.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Katna + LLaVA + LLMSampling Strategy=Katna key frame sampling, LMM=LLaVA-1.5, LLM=Vicuna-v1.5, Fusion=Late fusion2026.01 | 61.6 | — | — | — | — | — | — | — | 3.74 | — | — | — | — | |
| MM1.5-Video-7BTraining Mode=SFT, Video Data=true, Model Size=7B2024.09 | 60.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PLLaVALLM Size=34B, Vision Encoder=CLIP-L, Training Mode=Fine-tuned2024.07 | 60.9 | 3.7 | — | — | — | — | — | — | — | — | — | — | — | |
| ReMoRa2026.02 | 60.5 | — | — | — | — | — | — | — | 3.7 | — | — | — | — | |
| LLaVA-NeXT-Video-DPOLLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 60.2 | 3.5 | — | — | — | — | — | — | — | — | — | — | — | |
| LongVILALLM Size=7B2024.08 | 59.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SF-LLaVA-34BLLM Size=34B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 59.2 | 3.5 | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-NeXT-VideoLLM Size=34B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 58.8 | 3.4 | — | — | — | — | — | — | — | — | — | — | — | |
| CoPE-VideoLM-7BParameters=7B2026.02 | 58.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| IG-VLM (LLaVA v1.6)Vision Encoder=ViT-L, LLM Size=34B, Inference Vision=single, Inference LLM=single, Video Trained=X2024.03 | 58.4 | 3.5 | — | — | — | — | — | — | — | — | — | — | — | |
| IG-VLMLLM Size=34B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 58.4 | 3.5 | — | — | — | — | — | — | — | — | — | — | — | |
| Random Sampling + LLaVA + LLMSampling Strategy=Random Sampling, LMM=LLaVA-1.5, LLM=Vicuna-v1.5, Fusion=Late fusion2026.01 | 58.2 | — | — | — | — | — | — | — | 3.78 | — | — | — | — | |
| VILA-40BParameters=40B2026.02 | 58 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MM1.5-Video-3BTraining Mode=SFT, Video Data=true, Model Size=3B2024.09 | 57.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Gemini 1.5 ProZero-shot=true2024.07 | 57.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Gemini-1.5-Pro2024.08 | 57.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Gemini-1.5-ProInput Res.=768×768, Num Frames=322025.03 | 57.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| IG-VLM (CogAgent)Vision Encoder=CLIP-E, LLM Size=7B, Inference Vision=single, Inference LLM=single, Video Trained=X2024.03 | 57.3 | 3.6 | — | — | — | — | — | — | — | — | — | — | — | |
| IG-VLMModel Category=Video QA Methods2025.12 | 57.3 | — | — | — | — | — | — | — | 3.6 | — | — | — | — | |
| IG-VLM (LLaVA v1.6)Vision Encoder=ViT-L, LLM Size=13B, Inference Vision=single, Inference LLM=single, Video Trained=X2024.03 | 57.1 | 3.5 | — | — | — | — | — | — | — | — | — | — | — | |
| IG-VLM (GPT-4V)Vision Encoder=Unknown, LLM Size=GPT-4, Inference Vision=single, Inference LLM=single, Video Trained=X2024.03 | 57 | 3.5 | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4V2024.08 | 57 | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4VInput Res.=512×512, Num Frames=322025.03 | 57 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLAVA-OVLLM Size=7B2024.08 | 56.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVAOneVision-7BTraining Mode=SFT, Video Data=true, Model Size=7B2024.09 | 56.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-OV-7BParameters=7B2026.02 | 56.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-Video-7BParameters=7B2026.02 | 56.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PLLaVA-7BTraining Mode=SFT, Video Data=true, Model Size=7B2024.09 | 56.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PLLaVALLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Fine-tuned2024.07 | 56.3 | 3.5 | — | — | — | — | — | — | — | — | — | — | — | |
| Llama 3-V 70BZero-shot=true, Parameters=70B, Max frames=642024.07 | 56.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PLLAVALLM Size=7B2024.08 | 56.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| PLLaVAModel Category=Video QA Methods2025.12 | 56.3 | — | — | — | — | — | — | — | 3.5 | — | — | — | — | |
| VideoCocaEvaluation Protocol=Classification2023.11 | 56.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MM1.5-Video-1BTraining Mode=SFT, Video Data=true, Model Size=1B2024.09 | 56.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VideoCoCaZero-Shot=false, Video-level training (VT)=true2024.03 | 56.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-NeXT-ImageLLM Size=34B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 55.6 | 3.3 | — | — | — | — | — | — | — | — | — | — | — | |
| SlowFast-LLaVA-7BTraining Mode=Training-free, Video Data=false, Model Size=7B2024.09 | 55.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SF-LLaVA-7BLLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 55.5 | 3.4 | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-NeXT-Interleave-7BTraining Mode=SFT, Video Data=true, Model Size=7B2024.09 | 55.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Gemini-1.5-FlashInput Res.=768×768, Num Frames=322025.03 | 55.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Mobile-VideoGPT-1.5BTotal Params=1.6B, Input Res.=224×224, Num Frames=16, Throughput=41.02025.03 | 54.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| IG-VLM (LLaVA v1.6)Vision Encoder=ViT-L, LLM Size=7B, Inference Vision=single, Inference LLM=single, Video Trained=X2024.03 | 54.3 | 3.4 | — | — | — | — | — | — | — | — | — | — | — | |
| IG-VLM-7B (LLaVA-v1.6)Training Mode=Training-free, Video Data=false, Model Size=7B2024.09 | 54.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| IG-VLMLLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 54.3 | 3.4 | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL2.5-2BTotal Params=2.2B, Input Res.=448×448, Num Frames=32, Throughput=21.82025.03 | 54.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-NeXT-ImageLLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 53.8 | 3.2 | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-NeXT-VideoLLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 53.5 | 3.2 | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-NeXTBackbone=LLaVA-NeXT2026.01 | 53.3 | — | — | — | — | — | — | — | 3.43 | — | — | — | — | |
| SSCDBackbone=LLaVA-NeXT2026.01 | 53.3 | — | — | — | — | — | — | — | 3.31 | — | — | — | — | |
| Multiclips: Clip Sampling + Video-LLaVA + LLMSampling Strategy=Clip Sampling, LMM=Video-LLaVA, LLM=Vicuna-v1.5, Fusion=Late fusion2026.01 | 53.1 | — | — | — | — | — | — | — | 3.52 | — | — | — | — | |
| VideoLLaMA2.1LLM Size=7B2024.08 | 53 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-NeXT + Dino-HealBackbone=LLaVA-NeXT2026.01 | 53 | — | — | — | — | — | — | — | 3.38 | — | — | — | — | |
| IXC-2.5-7BParameters=7B2026.02 | 52.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Llama 3-V 8BZero-shot=true, Parameters=8B, Max frames=642024.07 | 52.7 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MM1.5-Video-7BTraining Mode=Training-free, Video Data=false, Model Size=7B2024.09 | 52.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-Mini-8BTotal Params=8.4B, Input Res.=336×336, Num Frames=1fps, Throughput=4.62025.03 | 52.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Gemini 1.0 UltraZero-shot=true2024.07 | 52.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| EMA-7BParameters=7B2026.02 | 52.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| EMA2026.02 | 52.1 | — | — | — | — | — | — | — | 3.5 | — | — | — | — | |
| Flash-VStreamLLM Size=7B2024.08 | 51.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Flash-VstreamModel Category=Video QA Methods2025.12 | 51.9 | — | — | — | — | — | — | — | 3.4 | — | — | — | — | |
| Mobile-VideoGPT-0.5BTotal Params=0.6B, Input Res.=224×224, Num Frames=16, Throughput=45.92025.03 | 51.6 | — | — | — | — | — | — | — | — | — | — | — | — | |
| FreeVALLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 51.2 | 3.5 | — | — | — | — | — | — | — | — | — | — | — | |
| Mirasol3BNumber of Frames=512, Combiner Type=Transformer, Evaluation Protocol=Open-ended2023.11 | 51.13 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MM1.5-Video-3BTraining Mode=Training-free, Video Data=false, Model Size=3B2024.09 | 50.9 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ST-LLMVision Size=1.3B, Modality=V, Pretrain Data=null, Finetune Data=400K2024.12 | 50.9 | — | — | — | — | — | — | — | 3.3 | — | — | — | — | |
| ShareGPT4VideoLLM Size=8B2024.08 | 50.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VideoGPT+LLM Size=3.8B, Vision Encoder=CLIP-L, Training Mode=Fine-tuned2024.07 | 50.6 | 3.6 | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVAOneVision-0.5BTraining Mode=SFT, Video Data=true, Model Size=0.5B2024.09 | 50.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-OneVision-0.5BTotal Params=1.0B, Input Res.=384×384, Num Frames=32, Throughput=22.72025.03 | 50.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Video-LLaMA2LLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Fine-tuned, Table Source=1(b)2024.07 | 50.3 | 3.4 | — | — | — | — | — | — | — | — | — | — | — | |
| Video-LLaMA2-7BTraining Mode=SFT, Video Data=true, Model Size=7B2024.09 | 50.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Video-LLaMA2LLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Fine-tuned2024.07 | 50.2 | 3.3 | — | — | — | — | — | — | — | — | — | — | — | |
| VideoLLaMA2LLM Size=7B2024.08 | 50.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| MiCo-Chat-7BLLM=Vicuna-7B, Res.=224, Zero-shot=true2024.06 | 50.1 | 3.3 | — | — | — | — | — | — | — | — | — | — | — | |
| Video-LaVIT2026.02 | 50.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Video-LaVIT2026.02 | 50.1 | — | — | — | — | — | — | — | 3.3 | — | — | — | — | |
| LongVA-7BParameters=7B2026.02 | 50 | — | — | — | — | — | — | — | — | — | — | — | — | |
| LongVA-7BTotal Params=7.4B, Input Res.=224×224, Num Frames=8, Throughput=9.22025.03 | 50 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Mirasol3BNumber of Frames=512, Combiner Type=TTM, Evaluation Protocol=Open-ended2023.11 | 49.85 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Gemini 1.0 ProZero-shot=true2024.07 | 49.8 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DeepStack-L-7BTraining Mode=Training-free, Video Data=false, Model Size=7B2024.09 | 49.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DeepStack-LLLM Size=7B, Vision Encoder=CLIP-L, Training Mode=Training-free2024.07 | 49.3 | 3.1 | — | — | — | — | — | — | — | — | — | — | — | |
| VideoChat2Vision Encoder=UMT-L, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=O2024.03 | 49.1 | 3.3 | — | — | — | — | — | — | — | — | — | — | — | |
| VideoChat2LLM=Vicuna-7B, Res.=224, Zero-shot=true2024.06 | 49.1 | 3.3 | — | — | — | — | — | — | — | — | — | — | — | |
| VideoChat2-7BTraining Mode=SFT, Video Data=true, Model Size=7B2024.09 | 49.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| VideoChat2LLM Size=7B, Vision Encoder=UMT-L, Training Mode=Fine-tuned2024.07 | 49.1 | 3.3 | — | — | — | — | — | — | — | — | — | — | — | |
| VideoChat2Vision Size=496M, Modality=V+I, Pretrain Data=37M, Finetune Data=2M2024.12 | 49.1 | — | — | — | — | — | — | — | 3.3 | — | — | — | — | |
| LLaVA-NeXT + TCDBackbone=LLaVA-NeXT2026.01 | 49.1 | — | — | — | — | — | — | — | 3.1 | — | — | — | — | |
| VideoChat2-7BTotal Params=7.3B, Input Res.=224×224, Num Frames=16, Throughput=11.42025.03 | 49.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Vista-LLaMAVision Encoder=CLIP-G, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=O2024.03 | 48.3 | 3.3 | — | — | — | — | — | — | — | — | — | — | — | |
| Vista-LLaMA-7BTraining Mode=SFT, Video Data=true, Model Size=7B2024.09 | 48.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Vista-LLaMALLM Size=7B, Vision Encoder=CLIP-G, Training Mode=Fine-tuned2024.07 | 48.3 | 3.3 | — | — | — | — | — | — | — | — | — | — | — |