Video Question Answering on VideoMME
85.1AccuracyGemini 2.5 Pro
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Gemini 2.5 Pro2026.01 | 85.1 | — | — | — | — | — | — | — | |
| Gemini-2.5-ProFrames=-2026.03 | 84.3 | — | — | — | — | — | — | — | |
| Gemini 1.5 ProFrames=1282026.01 | 78.6 | — | — | — | — | — | — | — | |
| Seed1.5-VLFrames=-2026.03 | 77.9 | — | — | — | — | — | — | — | |
| Gemini-1.5-FlashFrames=1282026.01 | 76.1 | — | — | — | — | — | — | — | |
| Gemini 1.5 Pro# Frames=-2024.10 | 75 | — | — | — | — | — | — | — | |
| Gemini-1.5-ProFrames=2562025.02 | 75 | — | — | — | — | — | — | — | |
| Gemini-1.5-Pro#Frames=1fps2026.02 | 75 | — | — | — | — | — | — | — | |
| Gemini-1.5-ProType=Proprietary, Model=–, Frames=2562026.05 | 75 | — | — | — | — | — | — | — | |
| Qwen3-VL-8BReasoning Mode=None, Reproduced=true2026.01 | 72.5 | — | — | — | — | — | 2.2 | — | |
| GPT 4o# Frames=-2024.10 | 71.9 | — | — | — | — | — | — | — | |
| GPT-4oFrames=2562025.02 | 71.9 | — | — | — | — | — | — | — | |
| GPT4-o#Frames=1fps2026.02 | 71.9 | — | — | — | — | — | — | — | |
| GPT-4oFrames=-2026.03 | 71.9 | — | — | — | — | — | — | — | |
| GPT-4o#params=−, #frames=−2026.04 | 71.9 | — | — | — | — | — | — | 60.02 | |
| GPT-4oType=Proprietary, Model=–, Frames=2562026.05 | 71.9 | — | — | — | — | — | — | — | |
| VideoAuto-R1Reasoning Mode=AutoThink, Backbone=Qwen3-VL-8B2026.01 | 71.7 | — | — | — | — | 11 | 52 | — | |
| Qwen2-VL-7B-InstructFrames=2fps2026.01 | 71.4 | — | — | — | — | — | — | — | |
| Gemini-1.5-FlashFrames=2562025.02 | 70.3 | — | — | — | — | — | — | — | |
| Gemini 1.5 Flash#params=−, #frames=−2026.04 | 70.3 | — | — | — | — | — | — | — | |
| Gemini-1.5-FlashType=Proprietary, Model=–, Frames=2562026.05 | 70.3 | — | — | — | — | — | — | — | |
| VanillaBackbone=Qwen3-VL-8B, Retention Ratio R=100%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 69.4 | — | — | — | — | — | — | — | |
| LLaVA-Video-7B+TIR-FlowFrames=322026.01 | 68.9 | — | — | — | — | — | — | — | |
| Qwen3-VL8B + OutRoModel=Qwen3-VL8B, Decoding=OutRo2026.03 | 68.33 | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7B+TIR-FlowFrames=322026.01 | 68.3 | — | — | — | — | — | — | — | |
| T*Type=Training-free, Model=LLaVA-OneVision-72B, Frames=322026.05 | 68.3 | — | — | — | — | — | — | — | |
| Qwen3-VL8BModel=Qwen3-VL8B2026.03 | 67.44 | — | — | — | — | — | — | — | |
| ResAdaptBackbone=Qwen3-VL-8B, Retention Ratio R=22.9%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 67.4 | — | — | — | — | — | — | — | |
| VideoAuto-R1Reasoning Mode=AutoThink, Backbone=Qwen2.5-VL-7B2026.01 | 67.3 | — | — | — | — | 40 | 44 | — | |
| Random DropBackbone=Qwen3-VL-8B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 67.2 | — | — | — | — | — | — | — | |
| ToMeBackbone=Qwen3-VL-8B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 67.2 | — | — | — | — | — | — | — | |
| VisionZipBackbone=Qwen3-VL-8B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 67.1 | — | — | — | — | — | — | — | |
| ResAdaptBackbone=Qwen3-VL-8B, Retention Ratio R=11.1%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 66.8 | — | — | — | — | — | — | — | |
| FixedScaleBackbone=Qwen3-VL-8B, Retention Ratio R=12.3%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 66.7 | — | — | — | — | — | — | — | |
| InternVL3-8BFrames=-2026.03 | 66.3 | — | — | — | — | — | — | — | |
| LOVE-R1Reasoning Mode=Think-Only2026.01 | 66.2 | — | — | — | — | — | — | — | |
| VideoLLaMA3-7BFrames=1802026.01 | 66.2 | — | — | — | — | — | — | — | |
| VideoLLaMA3-7BFrames=-2026.03 | 66.2 | — | — | — | — | — | — | — | |
| VideoAuto-R1 + ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=23.8%, Reasoning (CoT)=true, Frames=128 Frames2026.03 | 66.2 | — | — | — | — | — | — | — | |
| MIRAType=Training-free, Model=LLaVA-Video-7B, Frames=642026.05 | 66.2 | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7BReasoning Mode=None, Reproduced=true2026.01 | 66 | — | — | — | — | — | 3 | — | |
| InternVL3.5-8BFrames=-2026.03 | 66 | — | — | — | — | — | — | — | |
| ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=22.9%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 65.6 | — | — | — | — | — | — | — | |
| TPOType=Training-based, Model=LLaVA-Video-7B, Frames=642026.05 | 65.6 | — | — | — | — | — | — | — | |
| LongVULLM=Qwen2-7B, Frames=1fps2025.03 | 65.4 | — | — | — | — | — | — | — | |
| InternVL3.5-4BFrames=-2026.03 | 65.4 | — | — | — | — | — | — | — | |
| LLaVA-Video w/ AKSFrames=64, LLM=7B, sampling=AKS2025.02 | 65.3 | — | — | — | — | — | — | — | |
| Weaver#Frames=128+1fps2026.02 | 65.3 | — | — | — | — | — | — | — | |
| R-MSD (4B)Frames=642026.03 | 65.3 | — | — | — | — | — | — | — | |
| VanillaBackbone=Qwen2.5-VL-7B, Retention Ratio R=100%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 65.3 | — | — | — | — | — | — | — | |
| VideoChat-R1.5Reasoning Mode=Think-Only2026.01 | 65.2 | — | — | — | — | — | 133 | — | |
| TTA-Vidbase_model=InternVL-3, #params=8B, #frames=322026.04 | 65.11 | — | — | — | — | — | — | 55.13 | |
| Long-VILA-R1#Frames=5122026.02 | 65.1 | — | — | — | — | — | — | — | |
| LongVILA-R1Reasoning Mode=Think-Only2026.01 | 65.1 | — | — | — | — | — | — | — | |
| ToMeBackbone=Qwen2.5-VL-7B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 65.1 | — | — | — | — | — | — | — | |
| CATS (ours)†Type=Training-free, Model=LLaVA-Video-7B, Frames=322026.05 | 65.04 | — | — | — | — | — | — | — | |
| VanillaBackbone=Qwen3-VL-8B, Retention Ratio R=100%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 65 | — | — | — | — | — | — | — | |
| DToMAType=Training-free, Model=LLaVA-Video-7B, Frames=642026.05 | 65 | — | — | — | — | — | — | — | |
| Random DropBackbone=Qwen2.5-VL-7B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 64.9 | — | — | — | — | — | — | — | |
| VisionZipBackbone=Qwen2.5-VL-7B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 64.8 | — | — | — | — | — | — | — | |
| VideoChat-R1.5#params=7B, #frames=1282026.04 | 64.8 | — | — | — | — | — | — | 52.24 | |
| VideoAuto-R1Backbone=Qwen2.5-VL-7B, Retention Ratio R=100%, Reasoning (CoT)=true, Frames=128 Frames2026.03 | 64.7 | — | — | — | — | — | — | — | |
| VideoAuto-R1 + ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=11.4%, Reasoning (CoT)=true, Frames=128 Frames2026.03 | 64.7 | — | — | — | — | — | — | — | |
| ToMeBackbone=Qwen3-VL-8B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 64.7 | — | — | — | — | — | — | — | |
| BIMBAType=Training-based, Model=BIMBA-7B, Frames=1282026.05 | 64.7 | — | — | — | — | — | — | — | |
| BIMBA-LLaVALLM=Qwen2-7B, Frames=1282025.03 | 64.67 | — | — | — | — | — | — | — | |
| AKS†Type=Training-free, Model=LLaVA-Video-7B, Frames=322026.05 | 64.48 | — | — | — | — | — | — | — | |
| LLaVA-VideoFrames=64, LLM=7B2025.02 | 64.4 | — | — | — | — | — | — | — | |
| Video-R1#params=7B, #frames=1282026.04 | 64.3 | — | — | — | — | — | — | 51.94 | |
| InternVL2.5#Frames=16-642026.02 | 64.2 | — | — | — | — | — | — | — | |
| InternVL2.5-8BFrames=-2026.03 | 64.2 | — | — | — | — | — | — | — | |
| VisionZipBackbone=Qwen3-VL-8B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 64.2 | — | — | — | — | — | — | — | |
| VITALReasoning Mode=Think-Only2026.01 | 64.1 | — | — | — | — | — | — | — | |
| FixedScaleBackbone=Qwen2.5-VL-7B, Retention Ratio R=12.3%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 64.1 | — | — | — | — | — | — | — | |
| Random DropBackbone=Qwen3-VL-8B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 64.1 | — | — | — | — | — | — | — | |
| VideoChat-R1#params=7B, #frames=1282026.04 | 64.1 | — | — | — | — | — | — | 52.34 | |
| Video-RFT#params=7B, #frames=1282026.04 | 64.1 | — | — | — | — | — | — | 52.32 | |
| Original SFT+RL (4B)Frames=642026.03 | 64 | — | — | — | — | — | — | — | |
| Qwen3-VL-8B-InstructFrames=642026.03 | 64 | — | — | — | — | — | — | — | |
| FlashVidBackbone=Qwen3-VL-8B, Retention Ratio R=30.0%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 63.9 | — | — | — | — | — | — | — | |
| Weaver-SFT#Frames=128+1fps2026.02 | 63.8 | — | — | — | — | — | — | — | |
| Qwen3-VL-4B-InstructFrames=642026.03 | 63.8 | — | — | — | — | — | — | — | |
| ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=11.1%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 63.8 | — | — | — | — | — | — | — | |
| Video-R2#params=7B, #frames=1282026.04 | 63.8 | — | — | — | — | — | — | 53.92 | |
| TwigVLM++Tokens Retained Percentage=33.3%2025.03 | 63.6 | — | — | — | — | — | — | — | |
| Q2.5VL-7BTokens Retained Percentage=100%2025.03 | 63.4 | — | — | — | — | — | — | — | |
| LLaVA-VideoType=Foundational, Model=LLaVA-Video-7B, Frames=642026.05 | 63.33 | — | — | — | — | — | — | — | |
| Qwen2 VL (7B)# Frames=2 fps (max 768)2024.10 | 63.3 | — | — | — | — | — | — | — | |
| Qwen2-VLLLM=Qwen2-7B, Frames=2fps2025.03 | 63.3 | — | — | — | — | — | — | — | |
| LLaVA-VideoLLM=Qwen2-7B, Frames=642025.03 | 63.3 | — | — | — | — | — | — | — | |
| Qwen2.5-VL#Frames=1282026.02 | 63.3 | — | — | — | — | — | — | — | |
| LLaVA-Video-7BFrames=322026.01 | 63.3 | — | — | — | — | — | — | — | |
| LLaVA-VideoModel category=General Video LLMs2026.04 | 63.3 | — | — | — | — | — | — | — | |
| VideoAuto-R1Backbone=Qwen2.5-VL-7B, Retention Ratio R=100%, Reasoning (CoT)=true, Frames=32 Frames2026.03 | 63.2 | — | — | — | — | — | — | — | |
| Video-RTSReasoning Mode=Think-Only2026.01 | 63 | — | — | — | — | — | — | — | |
| Random DropBackbone=Qwen2.5-VL-7B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 63 | — | — | — | — | — | — | — | |
| Video-RTS#params=7B, #frames=1282026.04 | 63 | — | — | — | — | — | — | — | |
| STTMFrames=1fps2026.01 | 62.6 | — | — | — | — | — | — | — | |
| ResAdaptBackbone=Qwen3-VL-8B, Retention Ratio R=23.8%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 62.6 | — | — | — | — | — | — | — | |
| VideoLLaMA37B + OutRoModel=VideoLLaMA37B, Decoding=OutRo2026.03 | 62.56 | — | — | — | — | — | — | — |