Video Question-Answering on MLVU
76.2AccuracyVideoChat-A1
Evaluation Results
| Method | Links | |
|---|---|---|
| VideoChat-A1Backbone=InternVideo2.5-8B, Frames=41.3, Inference Time=28.4s2025.06 | 76.2 | |
| VideoLucyVenue=NeurIPS’25, Base Model=DeepSeek-R1, #Frames=-2025.10 | 76.1 | |
| VideoChat-A1Backbone=InternVL2.5-8B, Frames=35.9, Inference Time=14.7s2025.06 | 75.1 | |
| VideoChat-FlashSize=7B, #Tokens=162025.01 | 74.5 | |
| A.I.R.Venue=-, Base Model=InternVL3-8B, #Frames=≤322025.10 | 74.5 | |
| InternVideo2.5 (InternVL2.5+LRC)Size=7B, #Tokens=162025.01 | 72.8 | |
| InternVideo-2.5-8BFrames=512, Inference Time=33.8s2025.06 | 72.8 | |
| VanillaBackbone=Qwen3-VL-8B, Retention Ratio R=100%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 72.7 | |
| VideoRAGVenue=NeurIPS’25, Base Model=LLaVA-Video-7B, #Frames=642025.10 | 72.4 | |
| POINTS-LongNum Frame=248+8, Token/Frame=16, Total Num of Token=52482026.04 | 72.1 | |
| VideoChat-A1Backbone=Qwen2.5-VL-7B, Frames=42.0, Inference Time=18.6s2025.06 | 71.9 | |
| POINTS-LongNum Frame=248+8, Token/Frame=8, Total Num of Token=32002026.04 | 71.8 | |
| BIMBA-LLaVALLM=Qwen2-7B, Frames=1282025.03 | 71.37 | |
| POINTS-LongNum Frame=504+8, Token/Frame=16, Total Num of Token=93442026.04 | 71.2 | |
| LLaVA-VideoLLM=Qwen2-7B, Frames=642025.03 | 70.8 | |
| ResAdaptBackbone=Qwen3-VL-8B, Retention Ratio R=22.9%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 70.8 | |
| ToMeBackbone=Qwen3-VL-8B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 70.4 | |
| POINTS-LongNum Frame=504+8, Token/Frame=8, Total Num of Token=52482026.04 | 70.3 | |
| ResAdaptBackbone=Qwen3-VL-8B, Retention Ratio R=11.1%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 69.4 | |
| MemoryCard (MiniCPM-V-4.5)LLM Size=8B, #Frames=4+ 8+ 322026.06 | 69.4 | |
| A.I.R.Venue=-, Base Model=LLaVA-OV-7B, #Frames=≤322025.10 | 69.3 | |
| VisionZipBackbone=Qwen3-VL-8B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 69.2 | |
| InternVL2.5Size=7B, #Tokens=2562025.01 | 68.9 | |
| InternVL-2.5-8BFrames=64, Inference Time=12.4s2025.06 | 68.9 | |
| V-LynX-7BModel Category=Task-specific models ≥ 7B, ΔParams.=195.0M2026.05 | 68.4 | |
| FixedScaleBackbone=Qwen3-VL-8B, Retention Ratio R=12.3%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 67.7 | |
| Random DropBackbone=Qwen3-VL-8B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 67.4 | |
| ToMeBackbone=Qwen3-VL-8B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 67.3 | |
| PAVE-7B (w/ video feature)FLOPs(TB)=98.63, Total Params=8.2B, Trainable Params=170.5M2025.03 | 67 | |
| PAVE-7BModel Category=Task-specific models ≥ 7B, ΔParams.=500.5M2026.05 | 67 | |
| LLaVA-OneVision 32 frames + BOLTLLM Size=7B, Frames=322025.03 | 66.8 | |
| VisionZipBackbone=Qwen3-VL-8B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 66.8 | |
| Video Panels (Qwen-2.5VL 7B)#frames=180, Model Context Category=Long-context VLMs2025.09 | 66.8 | |
| Qwen-2.5VL 7B#frames=180, Model Context Category=Long-context VLMs2025.09 | 66.7 | |
| VanillaBackbone=Qwen2.5-VL-7B, Retention Ratio R=100%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 66.5 | |
| MemoryCard (Qwen3-VL)LLM Size=8B, #Frames=4+ 8+ 322026.06 | 66.5 | |
| Flash-VstreamTotal Num of Token=115202026.04 | 66.3 | |
| LLaVA-Video 7B#frames=64, Model Context Category=Medium-context VLMs2025.09 | 66.2 | |
| VideoAuto-R1 + ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=23.8%, Reasoning (CoT)=true, Frames=128 Frames2026.03 | 66 | |
| Video Panels (LLaVA-Video 7B)#frames=64, Model Context Category=Medium-context VLMs2025.09 | 66 | |
| ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=22.9%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 65.9 | |
| LLaVA-OneVision 16 frames + BOLTLLM Size=7B, Frames=162025.03 | 65.8 | |
| Video Panels (Qwen-2VL 7B)#frames=180, Model Context Category=Long-context VLMs2025.09 | 65.8 | |
| Qwen-2VL 7B#frames=180, Model Context Category=Long-context VLMs2025.09 | 65.7 | |
| MemoryCard (Qwen2-VL)LLM Size=7B, #Frames=4+ 8+ 322026.06 | 65.7 | |
| Frame-VoyagerLLM Size=8B2025.03 | 65.6 | |
| Frame-VoyagerLLM Size=8B, #Frames=82026.06 | 65.6 | |
| LongVUSize=7B, #Tokens=642025.01 | 65.4 | |
| LongVULLM Size=7B2025.03 | 65.4 | |
| LongVULLM Size=7B, #Frames=*2026.06 | 65.4 | |
| Video Panels (LLaVA-OV 7B)#frames=32, Model Context Category=Medium-context VLMs2025.09 | 65.3 | |
| VideoAuto-R1Backbone=Qwen2.5-VL-7B, Retention Ratio R=100%, Reasoning (CoT)=true, Frames=128 Frames2026.03 | 65.1 | |
| Video-XLLLM=Qwen2-7B, Frames=20482025.03 | 64.9 | |
| Video Panels (Qwen-2.5VL 7B)#frames=32, Model Context Category=Medium-context VLMs2025.09 | 64.9 | |
| Video-XLLLM Size=7B, #Frames=128/2562026.06 | 64.9 | |
| Random DropBackbone=Qwen2.5-VL-7B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 64.8 | |
| LLaVA-OneVisionSize=7B, #Tokens=1962025.01 | 64.7 | |
| LLaVA-OneVisionLLM=Qwen2-7B, Frames=322025.03 | 64.7 | |
| LLaVA-OV-7BFLOPs(TB)=98.53, Total Params=8.2B2025.03 | 64.7 | |
| POINTS1.5-8B-onlineNum Frame=64, Token/Frame=324, Total Num of Token=207362026.04 | 64.7 | |
| LLaVA-OV-7BModel Category=Task-specific models ≥ 7B2026.05 | 64.7 | |
| LLaVA-OneVisionLLM Size=7B, #Frames=*2026.06 | 64.7 | |
| GPT4-o2025.01 | 64.6 | |
| GPT-4oFrames=384, Inference Time=134.4s2025.06 | 64.6 | |
| VideoAuto-R1 + ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=11.4%, Reasoning (CoT)=true, Frames=128 Frames2026.03 | 64.6 | |
| VisionZipBackbone=Qwen2.5-VL-7B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 64.5 | |
| FixedScaleBackbone=Qwen2.5-VL-7B, Retention Ratio R=12.3%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 64.5 | |
| ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=11.1%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 64.3 | |
| VanillaBackbone=Qwen3-VL-8B, Retention Ratio R=100%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 64 | |
| ToMeBackbone=Qwen2.5-VL-7B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 63.5 | |
| LLaVA-OneVision 8 frames + BOLTLLM Size=7B, Frames=82025.03 | 63.4 | |
| Random DropBackbone=Qwen2.5-VL-7B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 63.4 | |
| LLaVA-OneVision 32 framesLLM Size=7B, Frames=322025.03 | 63.2 | |
| VisionZipBackbone=Qwen2.5-VL-7B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 63.2 | |
| VanillaBackbone=Qwen2.5-VL-7B, Retention Ratio R=100%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 63.1 | |
| ToMeBackbone=Qwen2.5-VL-7B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 63.1 | |
| Qwen2-VL-onlineTotal Num of Token=115202026.04 | 62.9 | |
| LLaVA-OV 7B#frames=32, Model Context Category=Medium-context VLMs2025.09 | 62.9 | |
| Random DropBackbone=Qwen3-VL-8B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 62.4 | |
| VanillaRet.=100%2026.05 | 62.1 | |
| FlashVidBackbone=Qwen3-VL-8B, Retention Ratio R=30.0%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 61.9 | |
| OTT-VidRet.=25%2026.05 | 61.8 | |
| FlashVIDRet.=25%2026.05 | 61.6 | |
| UniCompRet.=25%2026.05 | 61.6 | |
| HoliTomRet.=20%2026.05 | 61.6 | |
| OTT-VidRet.=20%2026.05 | 61.5 | |
| OTT-VidRet.=15%2026.05 | 61.4 | |
| HoliTomRet.=25%2026.05 | 61.3 | |
| UniCompRet.=20%2026.05 | 61.3 | |
| FlashVIDRet.=15%2026.05 | 61.3 | |
| LLaVA-OneVision 16 framesLLM Size=7B, Frames=162025.03 | 61.2 | |
| VideoAuto-R1 + ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=23.8%, Reasoning (CoT)=true, Frames=32 Frames2026.03 | 61.1 | |
| FastVIDRet.=25%2026.05 | 61.1 | |
| KangarooLLM=LLaMA2-8B, Frames=642025.03 | 61 | |
| ResAdaptBackbone=Qwen3-VL-8B, Retention Ratio R=23.8%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 61 | |
| HoliTomRet.=15%2026.05 | 61 | |
| ToMeBackbone=Qwen3-VL-8B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 60.8 | |
| FastVIDRet.=20%2026.05 | 60.8 | |
| FlashVIDRet.=20%2026.05 | 60.8 | |
| OTT-VidRet.=10%2026.05 | 60.7 |