Video Question Answering on MLVU (M-Avg)
76.2M-Avg ScoreQwen2-VL + ReQuest
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2-VL + ReQuestLLM Size=8B, #Frames=≤ 512, zero-shot transfer=true2026.07 | 76.2 | |
| Qwen2-VLLLM Size=8B, #Frames=512, reproduced=true2026.07 | 74 | |
| InternVL3 + ReFoCUSLLM Size=8B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 72.7 | |
| Qwen3-VL + ReFoCUSLLM Size=8B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 72.5 | |
| Video-RAGBase MLLM=LLaVA-Video 7B, Input Tokens=15k2026.03 | 72.4 | |
| FlexMemBase MLLM=LLaVA-Video 7B, Sampled Frames=512/1024frm, Input Tokens=13k2026.03 | 72.4 | |
| Video-o3 (SFT+RL)Size=7B2026.01 | 72.1 | |
| AKSBase MLLM=LLaVA-Video 7B, Sampled Frames=1fps, Input Tokens=13k2026.03 | 72 | |
| Video-o3 (RL)Size=7B2026.01 | 71.9 | |
| Qwen3-VL + ReFoCUSLLM Size=4B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 71.9 | |
| DToMABase MLLM=LLaVA-Video 7B, Input Tokens=12k2026.03 | 71.7 | |
| AdaRETAKEBase MLLM=LLaVA-Video 7B, Sampled Frames=1024frm, Input Tokens=40k2026.03 | 71.7 | |
| LLaVA-Video + ReQuestLLM Size=7B, #Frames=322026.07 | 71.7 | |
| InternVL3.5 + ReFoCUSLLM Size=4B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 71.5 | |
| LLaVA-Video 7BBase MLLM=LLaVA-Video 7B, Sampled Frames=64frm, Input Tokens=13k2026.03 | 71.2 | |
| LLaVA-Video + EFSParams=7B, Frames=642026.03 | 70.9 | |
| LLaVA-VideoSize=7B2026.01 | 70.8 | |
| InternVL3.5 + ReFoCUSLLM Size=8B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 70.6 | |
| Qwen2.5-VLSize=7B2026.01 | 70.2 | |
| NVILAParams=8B, Frames=2562026.03 | 70.1 | |
| NVILALLM Size=8B, #Frames=10242026.07 | 70.1 | |
| FlexMemBase MLLM=LLaVA-OV 7B, Sampled Frames=512/1024frm, Input Tokens=7k2026.03 | 68.9 | |
| LLaVa-OneVision + ReQuestLLM Size=7B, #Frames=32, zero-shot transfer=true2026.07 | 68.8 | |
| LLaVA-OneVision + ReFoCUSLLM Size=7B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 68.5 | |
| AKSBase MLLM=LLaVA-OV 7B, Sampled Frames=1fps, Input Tokens=7k2026.03 | 68.3 | |
| LLaVA-VideoParams=7B, Frames=642026.03 | 68.1 | |
| InternVL3LLM Size=8B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 68.1 | |
| InternVL3 + ReFoCUSLLM Size=2B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 68 | |
| LOVE-R1Size=7B2026.01 | 67.4 | |
| LLaVA-OneVision + EFSParams=7B, Frames=82026.03 | 67.4 | |
| InternVL3.5LLM Size=8B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 67.3 | |
| InternVL3.5LLM Size=4B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 66.6 | |
| Qwen2.5-VL + EFSParams=7B, Frames=162026.03 | 66 | |
| BOLTBase MLLM=LLaVA-OV 7B, Sampled Frames=1fps, Input Tokens=7k2026.03 | 65.8 | |
| Frame-VoyagerParams=7B, Frames=82026.03 | 65.6 | |
| Frame-VoyagerLLM Size=7B, #Frames=82026.07 | 65.6 | |
| LongVUParams=7B, Frames=1fps2026.03 | 65.4 | |
| LongVULLM Size=7B, #Frames=1fps2026.07 | 65.4 | |
| GPT-4o + ReFoCUSLLM Size=Closed, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 65.1 | |
| Video-XLParams=7B, Frames=128/2562026.03 | 64.9 | |
| LLaVA-VideoLLM Size=7B, #Frames=32, reproduced=true2026.07 | 64.7 | |
| GPT-4oParams=-, Frames=384/256/0.5fps2026.03 | 64.6 | |
| VideoMindSize=7B2025.03 | 64.4 | |
| AdaRETAKEBase MLLM=LLaVA-OV 7B, Sampled Frames=1024frm, Input Tokens=20k2026.03 | 64.4 | |
| LLaVA-OneVisionLLM Size=7B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 63.7 | |
| ConanSize=7B2026.01 | 63.4 | |
| LLaVA-OV 7BBase MLLM=LLaVA-OV 7B, Sampled Frames=32frm, Input Tokens=7k2026.03 | 63.4 | |
| Qwen3-VLLLM Size=4B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 63.1 | |
| LLaVa-OneVisionLLM Size=7B, #Frames=32, reproduced=true2026.07 | 63.1 | |
| Qwen3-VLLLM Size=8B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 63 | |
| InternVL3LLM Size=2B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 62.7 | |
| VideoTreeLLM Size=–, #Frames=–2026.07 | 60.4 | |
| VideoLLaMA 3 + ReFoCUSLLM Size=7B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 59.8 | |
| Video-MTRSize=7B2026.01 | 59.7 | |
| InternVL3 + ReFoCUSLLM Size=1B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 58.9 | |
| VideoMindSize=2B2025.03 | 58.7 | |
| GPT-4oLLM Size=Closed, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 58.7 | |
| LLaVA-OneVisionParams=7B, Frames=82026.03 | 58.6 | |
| Gemini 2.5 Flash + ReFoCUSLLM Size=Closed, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 58 | |
| Qwen2.5-VLParams=7B, Frames=162026.03 | 56.5 | |
| LongVASize=7B2025.03 | 56.3 | |
| LongVAParams=7B, Frames=128/2562026.03 | 56.3 | |
| LLoViLLM Size=–, #Frames=–2026.07 | 55.1 | |
| VideoChat-TPOSize=7B2025.03 | 54.7 | |
| GPT-4o2025.03 | 54.5 | |
| InternVL3LLM Size=1B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 54 | |
| PLLAVASize=34B2025.03 | 53.6 | |
| VideoLLaMA 3LLM Size=7B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 52.9 | |
| Gemini 2.5 FlashLLM Size=Closed, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 52.8 | |
| LLaVA-OneVision + ReFoCUSLLM Size=0.5B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 50.3 | |
| VideoLLaMA 3 + ReFoCUSLLM Size=2B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 50.2 | |
| Video-LLaVAParams=7B, Frames=82026.03 | 47.3 | |
| Video-LLaVALLM Size=7B, #Frames=82026.07 | 47.3 | |
| VideoLLaMA 3LLM Size=2B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 46.8 | |
| ShareGPT4VideoLLM Size=8B, #Frames=162026.07 | 46.4 | |
| LLaVA-OneVisionLLM Size=0.5B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 44.8 | |
| VideoChat2LLM Size=7B, #Frames=162026.07 | 44.5 | |
| TimeChatSize=7B2025.03 | 30.9 | |
| Video-LLaVASize=7B2025.03 | 29.3 | |
| MovieChatSize=7B2025.03 | 25.8 |