Video Question-Answering on EgoSchema (test)
77.9AccuracyQwenVL2
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| QwenVL2Size=72B2025.01 | 77.9 | — | |
| GPT4-o2025.01 | 72.2 | — | |
| Gemini-1.5-Pro2024.08 | 72.2 | — | |
| Gemini-1.5-Pro2025.01 | 71.2 | — | |
| VideoChat-TLLM Size=7B2024.10 | 68.4 | — | |
| Vgent(M)LLM Backbone=Qwen2-VL-7B, Input Processing=Video-to-Text Tools2026.04 | 68 | — | |
| LongVILASize=7B, #Tokens=1962025.01 | 67.7 | — | |
| LongVILALLM Size=7B2024.08 | 67.7 | — | |
| LongVUSize=7B, #Tokens=642025.01 | 67.6 | — | |
| VideoStir(M)LLM Backbone=Qwen2-VL-7B, Input Processing=Native Input2026.04 | 67.2 | — | |
| QwenVL2Size=7B2025.01 | 66.7 | — | |
| DrVideo(M)LLM Backbone=GPT-4, Input Processing=Video-to-Text Tools2026.04 | 66.4 | — | |
| VideoTreeLLM Size=GPT-4, Vision Encoder=Unknown, Training Strategy=SFT2024.07 | 66.2 | — | |
| VideoTree(M)LLM Backbone=GPT-4, Input Processing=Video-to-Text Tools2026.04 | 66.2 | — | |
| IG-VLM(M)LLM Backbone=Qwen2-VL-7B, Input Processing=Native Input2026.04 | 66.2 | — | |
| VideoLLaMA2Size=72B, #Tokens=722025.01 | 63.9 | — | |
| InternVideo2.5 (InternVL2.5+LRC)Size=7B, #Tokens=162025.01 | 63.9 | — | |
| VideoChat2LLM Size=7B2024.10 | 63.6 | — | |
| VideoAgentLLM Size=GPT-4, Reference=Fan et al., 20242024.10 | 62.8 | — | |
| VidAgent(M)LLM Backbone=GPT-4, Input Processing=Video-to-Text Tools2026.04 | 62.8 | — | |
| KangarooLLM Size=8B2024.08 | 62.7 | — | |
| DrVideo(M)LLM Backbone=GPT-3.5, Input Processing=Video-to-Text Tools2026.04 | 62.6 | — | |
| LLaVA-OneVisionSize=72B, #Tokens=1962025.01 | 61.3 | — | |
| LLoVi(M)LLM Backbone=GPT-4, Input Processing=Video-to-Text Tools2026.04 | 61.2 | — | |
| Vista-LLAMAVision Encoder=CLIP-G, LLM Size=7B, Inference Vision=multiple, Inference LLM=single, Video Trained=true2024.03 | 60.7 | — | |
| VideoAgentLLM Size=GPT-4, Vision Encoder=Unknown, Training Strategy=SFT2024.07 | 60.2 | — | |
| VideoAgentLLM Size=GPT-4, Reference=Wang et al., 2024c2024.10 | 60.2 | — | |
| VideoAgent(M)LLM Backbone=GPT-4, Input Processing=Video-to-Text Tools2026.04 | 60.2 | — | |
| LLaVA-OneVisionSize=7B, #Tokens=1962025.01 | 60.1 | — | |
| LLAVA-OVLLM Size=7B2024.08 | 60.1 | — | |
| InternVideo2-HDSize=7B, #Tokens=722025.01 | 60 | — | |
| GPT-4VVision Encoder=Unknown, LLM Size=GPT-4, Inference Vision=single, Inference LLM=single, Video Trained=false2024.03 | 59.8 | — | |
| LLoVi(M)LLM Backbone=GPT-3.5, Input Processing=Video-to-Text Tools2026.04 | 57.6 | — | |
| ProViQVision Encoder=TimeSformer-L, LLM Size=GPT-3.5, Inference Vision=multiple, Inference LLM=multiple, Video Trained=false2024.03 | 57.1 | — | |
| ProViQZero-shot=true2023.12 | 57.1 | — | |
| SF-LLaVA-34BLLM Size=34B, Vision Encoder=CLIP-L, Training Strategy=Training-free2024.07 | 55.8 | — | |
| LLaVA v1.6Vision Encoder=ViT-L, LLM Size=34B, Inference Vision=single, Inference LLM=single, Video Trained=false2024.03 | 53.6 | — | |
| IG-VLM (LLAVA-v1.6)LLM Size=34B, Vision Encoder=CLIP-L, Training Strategy=Training-free2024.07 | 53.6 | — | |
| VideoLLaMA2.1LLM Size=7B2024.08 | 53.1 | — | |
| mPLUG-Owl3Size=7B2025.01 | 52.1 | — | |
| VideoLLaMA2Size=7B, #Tokens=722025.01 | 51.7 | — | |
| VideoLLaMA2LLM Size=7B2024.08 | 51.7 | — | |
| MoReVQAFT (Fine-tuned)=false, Number of video frames (n)=302024.04 | 51.7 | — | |
| InternVL2.5Size=7B, #Tokens=2562025.01 | 51.5 | — | |
| LLoViVision Encoder=ViT-L, LLM Size=GPT-3.5, Inference Vision=multiple, Inference LLM=multiple, Video Trained=false2024.03 | 50.3 | — | |
| LLoViLLM Size=GPT-3.5, Vision Encoder=Unknown, Training Strategy=SFT2024.07 | 50.3 | — | |
| JCEFFT (Fine-tuned)=false2024.04 | 50 | — | |
| IG-VLM LLaVA v1.6Vision Encoder=ViT-L, LLM Size=7B, Inference Vision=single, Inference LLM=single, Video Trained=false2024.03 | 47 | — | |
| MC-ViT-LParams=424M, Evaluation Protocol=finetuned2024.05 | 44.4 | — | |
| VideoStreamingParams=7B+1.3B, Evaluation Protocol=zero-shot2024.05 | 44.1 | — | |
| VanillaBase Model=LLaVA-NeXT-Video-7B, Retained Tokens=Full, Pruning Rate=0%2026.01 | 41.4 | 100 | |
| LangRepoParams=8x7B, Evaluation Protocol=zero-shot2024.05 | 41.2 | — | |
| SparseVLMBase Model=LLaVA-NeXT-Video-7B, Retained Tokens=128, Pruning Rate=88.9%2026.01 | 40.6 | 98 | |
| ViTCoPBase Model=LLaVA-NeXT-Video-7B, Retained Tokens=128, Pruning Rate=88.9%2026.01 | 40.5 | 97.7 | |
| LangRepoParams=7B, Evaluation Protocol=zero-shot2024.05 | 38.9 | — | |
| TimeMambazero-shot=true, #Frame=8192, training dataset=Ego4D, training frames=42024.03 | 38.7 | — | |
| Video-LLaVALLM Size=7B2024.08 | 38.4 | — | |
| VisionZipBase Model=LLaVA-NeXT-Video-7B, Retained Tokens=128, Pruning Rate=88.9%2026.01 | 37 | 89.3 | |
| VamosParams=13B, Evaluation Protocol=zero-shot2024.05 | 36.7 | — | |
| LLaVA v1.6Vision Encoder=ViT-L, LLM Size=7B, Inference Vision=single, Inference LLM=single, Video Trained=false2024.03 | 35.8 | — | |
| PyramidDropBase Model=LLaVA-NeXT-Video-7B, Retained Tokens=128, Pruning Rate=88.9%2026.01 | 35.7 | 86.3 | |
| FastVBase Model=LLaVA-NeXT-Video-7B, Retained Tokens=128, Pruning Rate=88.9%2026.01 | 34.5 | 83.2 | |
| TimeSformerzero-shot=true, #Frame=16, training dataset=Ego4D, training frames=42024.03 | 33.6 | — | |
| TimeSformerzero-shot=true, #Frame=8192, training dataset=Ego4D, training frames=42024.03 | 33.5 | — | |
| LLoViParams=7B, Evaluation Protocol=zero-shot2024.05 | 33.5 | — | |
| TimeMambazero-shot=true, #Frame=128, training dataset=Ego4D, training frames=42024.03 | 33.4 | — | |
| LongViViTParams=1B, Evaluation Protocol=finetuned2024.05 | 33.3 | — | |
| LongViViTFT (Fine-tuned)=false, Concurrent=true2024.04 | 33.3 | — | |
| TimeSformerzero-shot=true, #Frame=128, training dataset=Ego4D, training frames=42024.03 | 32.9 | — | |
| InternVideoVision Encoder=ViT-L, LLM Size=1.3B, Inference Vision=multiple, Inference LLM=single, Video Trained=true2024.03 | 32.1 | — | |
| InternVideoParams=478M, Evaluation Protocol=zero-shot2024.05 | 32.1 | — | |
| InternVideoFT (Fine-tuned)=true2024.04 | 32.1 | — | |
| InternVideoZero-shot=true2023.12 | 32.1 | — | |
| InternVideozero-shot=true, #Frame=902024.03 | 32 | — | |
| TimeMambazero-shot=true, #Frame=16, training dataset=Ego4D, training frames=42024.03 | 31.9 | — | |
| InternVideozero-shot=true, #Frame=152024.03 | 31.6 | — | |
| mPLUG-OwlFT (Fine-tuned)=false2024.04 | 31.1 | — | |
| mPLUG-OWLZero-shot=true2023.12 | 31.1 | — | |
| ShortViViTFT (Fine-tuned)=false, Concurrent=true2024.04 | 31 | — | |
| FrozenBiLMzero-shot=true, #Frame=902024.03 | 26.9 | — | |
| FrozenBiLMVision Encoder=ViT-L, LLM Size=1B, Inference Vision=multiple, Inference LLM=single, Video Trained=true2024.03 | 26.9 | — | |
| FrozenBiLMParams=890M, Evaluation Protocol=zero-shot2024.05 | 26.9 | — | |
| FrozenBiLMFT (Fine-tuned)=false2024.04 | 26.9 | — | |
| FrozenBiLMZero-shot=true2023.12 | 26.9 | — | |
| FrozenBiLMzero-shot=true, #Frame=102024.03 | 26.4 | — | |
| CogAgentVision Encoder=CLIP-E, LLM Size=7B, Inference Vision=single, Inference LLM=single, Video Trained=false2024.03 | 25 | — | |
| SeViLAParams=4B, Evaluation Protocol=zero-shot2024.05 | 22.7 | — | |
| SeViLAFT (Fine-tuned)=false2024.04 | 22.7 | — | |
| VIOLETFT (Fine-tuned)=false2024.04 | 19.9 | — | |
| VIOLETZero-shot=true2023.12 | 19.9 | — |