Video Question Answering on VideoMMMU
74.9AccuracyGemini 2.5 Pro
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Gemini 2.5 Pro2026.01 | 74.9 | — | — | |
| Qwen3-VLModel Category=Video MLLMs, Release Date=Oct 252026.04 | 68.7 | — | — | |
| LFS + Qwen3-VL-8BZero-shot=true2026.01 | 66.8 | — | — | |
| Qwen3-VL-8BZero-shot=true2026.01 | 65.3 | — | — | |
| Qwen2-VL-7B-InstructFrames=2fps2026.01 | 65.3 | — | — | |
| VideoAuto-R1Reasoning Mode=AutoThink, Backbone=Qwen3-VL-8B2026.01 | 65 | 53 | 52 | |
| Qwen3-OmniModel Category=Omnimodal MLLMs, Release Date=Sep 252026.04 | 64.1 | — | — | |
| Gemini2.5-Flash-LiteZero-shot=true2026.01 | 63 | — | — | |
| VanillaBackbone=Qwen3-VL-8B, Retention Ratio R=100%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 63 | — | — | |
| GPT-4oLLM Size=Closed, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 62.9 | — | — | |
| GPT-4o + ReFoCUSLLM Size=Closed, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 62.1 | — | — | |
| GPT-4o#params=−, #frames=−2026.04 | 61.22 | — | — | |
| GPT4-o#Frames=1fps2026.02 | 61.2 | — | — | |
| GPT-4oFrames=642026.01 | 61.2 | — | — | |
| GPT-4oModel Category=Closed-Source APIs, Release Date=May 242026.04 | 61.2 | — | — | |
| Qwen3-VL + ReFoCUSLLM Size=8B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 61.1 | — | — | |
| Qwen3-VL-8BReasoning Mode=None, Reproduced=true2026.01 | 61 | — | 2.2 | |
| Qwen3-VL-8B + SynRLModel Scale=8B, SynRL Training=true2026.03 | 61 | — | — | |
| VanillaBackbone=Qwen3-VL-8B, Retention Ratio R=100%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 60.8 | — | — | |
| Qwen2.5-VL-72BZero-shot=true2026.01 | 60.2 | — | — | |
| Qwen3-VL-8BModel Scale=8B, SynRL Training=false2026.03 | 60 | — | — | |
| Omni-o3Model Category=Omnimodal MLLMs2026.04 | 59.9 | — | — | |
| ResAdaptBackbone=Qwen3-VL-8B, Retention Ratio R=22.9%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 59.6 | — | — | |
| Qwen3-VLLLM Size=8B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 59.1 | — | — | |
| VideoAuto-R1Reasoning Mode=AutoThink, Backbone=Qwen2.5-VL-7B2026.01 | 58.6 | 51 | 44 | |
| Qwen2.5-VL-7B+TIR-FlowFrames=322026.01 | 58.4 | — | — | |
| ResAdaptBackbone=Qwen3-VL-8B, Retention Ratio R=23.8%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 58.4 | — | — | |
| ResAdaptBackbone=Qwen3-VL-8B, Retention Ratio R=11.1%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 58.2 | — | — | |
| LLaVA-onevision-7BFrames=322026.01 | 57.1 | — | — | |
| VisionZipBackbone=Qwen3-VL-8B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 56.8 | — | — | |
| Qwen3-VL + ReFoCUSLLM Size=4B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 56.4 | — | — | |
| FixedScaleBackbone=Qwen3-VL-8B, Retention Ratio R=12.3%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 56.3 | — | — | |
| Qwen3-VL-4BZero-shot=true2026.01 | 56.2 | — | — | |
| ResAdaptBackbone=Qwen3-VL-8B, Retention Ratio R=11.4%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 56.1 | — | — | |
| Random DropBackbone=Qwen3-VL-8B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 55.5 | — | — | |
| Random DropBackbone=Qwen3-VL-8B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 55.3 | — | — | |
| FlashVidBackbone=Qwen3-VL-8B, Retention Ratio R=30.0%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 55.1 | — | — | |
| Qwen2.5-VL-7BReasoning Mode=None, Reproduced=true2026.01 | 54.7 | — | 3 | |
| Qwen3-VL-4B + SynRLModel Scale=4B, SynRL Training=true2026.03 | 54.5 | — | — | |
| VITALReasoning Mode=Think-Only2026.01 | 54.2 | — | — | |
| Qwen3-VL-4BModel Scale=4B, SynRL Training=false2026.03 | 54.1 | — | — | |
| Qwen3-VLLLM Size=4B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 54 | — | — | |
| Gemini-1.5-Pro#Frames=1fps2026.02 | 53.9 | — | — | |
| Gemini 1.5 ProFrames=1282026.01 | 53.9 | — | — | |
| Gemini1.5ProModel Category=Closed-Source APIs, Release Date=Dec 242026.04 | 53.9 | — | — | |
| AVATARModel Category=Omnimodal MLLMs, Release Date=CVPR 262026.04 | 53.7 | — | — | |
| TTA-Vidbase_model=InternVL-3, #params=8B, #frames=322026.04 | 53.66 | — | — | |
| VideoAuto-R1Backbone=Qwen2.5-VL-7B, Retention Ratio R=100%, Reasoning (CoT)=true, Frames=32 Frames2026.03 | 53.6 | — | — | |
| ToMeBackbone=Qwen3-VL-8B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 53.5 | — | — | |
| Random DropBackbone=Qwen3-VL-8B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 53.4 | — | — | |
| InternVL3.5 + ReFoCUSLLM Size=4B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 53.3 | — | — | |
| Qwen2.5-OmniModel Category=Omnimodal MLLMs, Release Date=Mar 252026.04 | 53.2 | — | — | |
| InternVL3.5 + ReFoCUSLLM Size=8B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 53.2 | — | — | |
| Video-RTSSize=7B2026.01 | 52.7 | — | — | |
| Video-RTSReasoning Mode=Think-Only2026.01 | 52.7 | — | — | |
| Video-RTS#params=7B, #frames=1282026.04 | 52.7 | — | — | |
| VideoAuto-R1 + ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=23.8%, Reasoning (CoT)=true, Frames=128 Frames2026.03 | 52.6 | — | — | |
| FixedScaleBackbone=Qwen3-VL-8B, Retention Ratio R=12.3%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 52.6 | — | — | |
| Video-R1#Frames=642026.02 | 52.4 | — | — | |
| Video-R1Size=7B2026.01 | 52.4 | — | — | |
| Video-R1-7BFrames=322026.01 | 52.3 | — | — | |
| Open-o3-videoModel Category=Video MLLMs, Release Date=Oct 252026.04 | 52.3 | — | — | |
| Video-ZoomerSize=7B2026.01 | 52.2 | — | — | |
| VideoAuto-R1Backbone=Qwen2.5-VL-7B, Retention Ratio R=100%, Reasoning (CoT)=true, Frames=128 Frames2026.03 | 52.2 | — | — | |
| VideoChat-R1#params=7B, #frames=1282026.04 | 52 | — | — | |
| InternVL3.5LLM Size=4B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 52 | — | — | |
| Rewatch-R1Size=7B2026.01 | 51.9 | — | — | |
| Video-o3 (SFT+RL)Size=7B2026.01 | 51.7 | — | — | |
| VisionZipBackbone=Qwen3-VL-8B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 51.5 | — | — | |
| Video-R1Reasoning Mode=Think-Only2026.01 | 51.4 | — | 386 | |
| VideoRFTModel Category=Video MLLMs, Release Date=May 252026.04 | 51.4 | — | — | |
| Weaver#Frames=128+1fps2026.02 | 51.3 | — | — | |
| VideoAuto-R1 + ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=11.4%, Reasoning (CoT)=true, Frames=128 Frames2026.03 | 51.3 | — | — | |
| Qwen2.5-VL-7BFrames=322026.01 | 51.2 | — | — | |
| VideoAuto-R1 + ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=23.8%, Reasoning (CoT)=true, Frames=32 Frames2026.03 | 51.2 | — | — | |
| VideoRFT#Frames=322026.02 | 51.1 | — | — | |
| Video-RFTReasoning Mode=Think-Only2026.01 | 51.1 | — | — | |
| ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=22.9%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 51.1 | — | — | |
| VideoChat-R1Model Category=Video MLLMs, Release Date=Apr 252026.04 | 51.1 | — | — | |
| HumanOmniV2Model Category=Omnimodal MLLMs, Release Date=June 252026.04 | 51.1 | — | — | |
| Video-R2#params=7B, #frames=1282026.04 | 50.8 | — | — | |
| InternVL3 + ReFoCUSLLM Size=8B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 50.6 | — | — | |
| VideoChat-R1.5#params=7B, #frames=1282026.04 | 50 | — | — | |
| Video-o3 (RL)Size=7B2026.01 | 50 | — | — | |
| InternVL3.5LLM Size=8B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 50 | — | — | |
| Gemini-1.5-FlashFrames=1282026.01 | 49.8 | — | — | |
| Gemini 1.5 Flash#params=−, #frames=−2026.04 | 49.78 | — | — | |
| VideoChat-R1.5Reasoning Mode=Think-Only2026.01 | 49.6 | — | 133 | |
| VanillaBackbone=Qwen2.5-VL-7B, Retention Ratio R=100%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 49.6 | — | — | |
| TTA-Vidbase_model=Qwen2.5-VL, #params=7B, #frames=322026.04 | 49.44 | — | — | |
| InternVL-3#params=8B, #frames=322026.04 | 49.33 | — | — | |
| InternVL3LLM Size=8B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 49.3 | — | — | |
| ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=11.1%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 49.2 | — | — | |
| VideoAuto-R1 + ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=11.4%, Reasoning (CoT)=true, Frames=32 Frames2026.03 | 49.1 | — | — | |
| ToMeBackbone=Qwen3-VL-8B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 49.1 | — | — | |
| Weaver-SFT#Frames=128+1fps2026.02 | 48.8 | — | — | |
| ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=23.8%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 48.8 | — | — | |
| Video-RFT#params=7B, #frames=1282026.04 | 48.1 | — | — | |
| Random DropBackbone=Qwen2.5-VL-7B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 48.1 | — | — | |
| VanillaBackbone=Qwen2.5-VL-7B, Retention Ratio R=100%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 47.9 | — | — |