Long Video Understanding on MLVU (dev)
78.1ScoreQwen2.5-VL+AdaRETAKE
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen2.5-VL+AdaRETAKELLM Size=72B2025.03 | 78.1 | — | |
| MM-MemBase Model=Qwen3-VL-8B, Category=Agent-based Systems2026.03 | 77.2 | — | |
| InternVL2.5LLM Size=72B2025.03 | 75.7 | — | |
| Qwen2.5-VL+AdaRETAKELLM Size=7B2025.03 | 75 | — | |
| Qwen2.5-VLLLM Size=72B2025.03 | 74.6 | — | |
| LLaVA-VideoLLM Size=72B2025.03 | 74.4 | — | |
| LLaVA-Video-72BCategory=Open-Sourced MLLMs2026.03 | 73.1 | — | |
| VideoLLaMA3LLM Size=7B2025.03 | 73 | — | |
| VideoLLaMA 3-7BCategory=Open-Sourced MLLMs2026.03 | 73 | — | |
| VideoRAGCategory=Agent-based Systems2026.03 | 72.4 | — | |
| Oryx-1.5LLM Size=32B2025.03 | 72.3 | — | |
| VgentCategory=Agent-based Systems2026.03 | 72.1 | — | |
| Qwen2-VL+AdaRETAKELLM Size=7B2025.03 | 72 | — | |
| TPOLLM Size=7B2025.03 | 71.1 | — | |
| GLM-4V-Plus2025.03 | 70.8 | — | |
| LLaVA-Video+AdaRETAKELLM Size=7B2025.03 | 70.6 | — | |
| AriaLLM Size=8x3.5B2025.03 | 70.6 | — | |
| Qwen2.5-VLLLM Size=7B2025.03 | 70.2 | — | |
| NVILALLM Size=8B2025.03 | 70.1 | — | |
| ByteVideoLLMLLM Size=14B2025.03 | 70.1 | — | |
| AdaSparkModel Size=Mid Size, Backbone=Qwen-2.5-VL-7B, SFT=true2026.04 | 69.8 | — | |
| VideoZoomerSize=7B, Evaluation Protocol=Max 128 frames2025.12 | 68.8 | — | |
| Qwen-2.5-VL-7BModel Size=Mid Size, SFT=false2026.04 | 68.3 | — | |
| Qwen-2.5-VL-7B + SFTModel Size=Mid Size, SFT=true2026.04 | 68.1 | — | |
| LLaVA-OneVisionLLM Size=72B2025.03 | 68 | — | |
| AdaSparkModel Size=Small Size, Backbone=Qwen-2.5-VL-3B, SFT=true2026.04 | 67.3 | — | |
| LLaVA-VideoLLM Size=7B2025.03 | 67 | — | |
| Qwen2-VLLLM Size=7B2025.03 | 66.9 | — | |
| FrameFusionModel Size=Mid Size, Backbone=Qwen-2.5-VL-7B, SFT=true2026.04 | 66.3 | — | |
| ToMeModel Size=Mid Size, Backbone=Qwen-2.5-VL-7B, SFT=true2026.04 | 66.1 | — | |
| Qwen3-VL-8BCategory=Open-Sourced MLLMs2026.03 | 65.9 | — | |
| VideoChat-Flash-2BModel Size=Small Size2026.04 | 65.7 | — | |
| FastVModel Size=Mid Size, Backbone=Qwen-2.5-VL-7B, SFT=true2026.04 | 65.7 | — | |
| LongVUSize=7B, Evaluation Protocol=Standard2025.12 | 65.4 | — | |
| ToMeModel Size=Small Size, Backbone=Qwen-2.5-VL-3B, SFT=true2026.04 | 65.4 | — | |
| Qwen-2.5-VL-3BModel Size=Small Size, SFT=false2026.04 | 65.3 | — | |
| Qwen-2.5-VL-3B + SFTModel Size=Small Size, SFT=true2026.04 | 65.2 | — | |
| VideoMinerCategory=Agent-based Systems2026.03 | 65.1 | — | |
| FastVModel Size=Small Size, Backbone=Qwen-2.5-VL-3B, SFT=true2026.04 | 65.1 | — | |
| Video-R1Size=7B, Evaluation Protocol=Max 128 frames2025.12 | 65 | — | |
| FrameFusionModel Size=Small Size, Backbone=Qwen-2.5-VL-3B, SFT=true2026.04 | 65 | — | |
| Video-XLSize=7B, Evaluation Protocol=Standard2025.12 | 64.9 | — | |
| Video-XL-7BModel Size=Mid Size2026.04 | 64.9 | — | |
| LLaVA-OneVisionSize=7B, Evaluation Protocol=Standard2025.12 | 64.7 | — | |
| MoBAModel Size=Mid Size, Backbone=Qwen-2.5-VL-7B, SFT=true2026.04 | 64.7 | — | |
| GPT-4o2025.03 | 64.6 | — | |
| GPT-4oSize=-, Evaluation Protocol=Standard2025.12 | 64.6 | — | |
| GPT-4oCategory=Proprietary Models2026.03 | 64.6 | — | |
| mPLUG-Owl3LLM Size=7B2025.03 | 63.7 | — | |
| MoBAModel Size=Small Size, Backbone=Qwen-2.5-VL-3B, SFT=true2026.04 | 63.2 | — | |
| InternVL2.5-2BModel Size=Small Size2026.04 | 61.4 | — | |
| KangarooSize=8B, Evaluation Protocol=Standard2025.12 | 61 | — | |
| VITA 1.5-7BCategory=Open-Sourced MLLMs2026.03 | 60.4 | — | |
| VideoTreeCategory=Agent-based Systems2026.03 | 60.4 | — | |
| Qwen2.5-VLSize=7B, Evaluation Protocol=Max 128 frames2025.12 | 58.3 | — | |
| VILA-1.5Size=7B, Evaluation Protocol=Standard2025.12 | 56.7 | — | |
| LongVASize=7B, Evaluation Protocol=Standard2025.12 | 56.3 | — | |
| LongVA-7BModel Size=Mid Size2026.04 | 56.3 | — | |
| LongVU-3BModel Size=Small Size2026.04 | 55.9 | — | |
| GPT-4VCategory=Proprietary Models2026.03 | 49.2 | — | |
| VideoChat2-7BModel Size=Mid Size2026.04 | 47.9 | — | |
| Video-LLaVASize=7B, Evaluation Protocol=Standard2025.12 | 36.2 | — | |
| LLaMA-VID-7BModel Size=Mid Size2026.04 | 33.2 | — | |
| Dispider-7BFr. (Sampling Rate/Frames)=1fps, Design Category=Streaming-Designed, Training Protocol=Training-based2025.10 | — | 61.7 | |
| Dream-VL#S (Model Size)=7B, #F (Number of input frames)=-, Architecture=DLM Baseline2026.01 | — | 61.1 | |
| DynFocusSize (LLM)=7B (Vicuna-v1.5), #Visual Tokens=322026.03 | — | 49.6 | |
| Flash-VStreamSize (LLM)=7B (Qwen2), #Visual Tokens=1282026.03 | — | 66.3 | |
| Frame-VoyagerSize (LLM)=7B (Qwen2), #Visual Tokens=1962026.03 | — | 65.6 | |
| Gemini-2.0-FlashModel Variant=Flash2026.02 | — | 71 | |
| Gemini-2.5-ProModel Variant=Pro2026.02 | — | 81.2 | |
| GPT-4o2026.02 | — | 64.6 | |
| GPT-4oSize (LLM)=-, #Visual Tokens=-2026.03 | — | 64.6 | |
| GPT-4oFrames=-, LLM Size=-2026.06 | — | 64.6 | |
| GPT-4VFrames=-, LLM Size=-2026.06 | — | 49.2 | |
| InternVideo2.5-8BBackbone=8B2026.02 | — | 72.8 | |
| InternVL2.5#S (Model Size)=7B, #F (Number of input frames)=64, Architecture=AR Baseline2026.01 | — | 68.9 | |
| LLaDA-V#S (Model Size)=8B, #F (Number of input frames)=32, Architecture=DLM Baseline, reproduced=true2026.01 | — | 59.4 | |
| LLaMA-VIDSize (LLM)=7B (Vicuna-v1.5), #Visual Tokens=22026.03 | — | 33.2 | |
| LLaVA-MiniSize (LLM)=7B (Vicuna-v1.5), #Visual Tokens=12026.03 | — | 42.8 | |
| LLaVA-OneVision#S (Model Size)=7B, #F (Number of input frames)=32, Architecture=AR Baseline, reproduced=true2026.01 | — | 64.7 | |
| LLaVA-OneVision-1.5Frames=32, LLM Size=8B2026.06 | — | 60.1 | |
| LLaVA-OneVision-1.5 w/ Q-FoldFrames=32, LLM Size=8B2026.06 | — | 64.6 | |
| LLaVA-OV-7BFr. (Sampling Rate/Frames)=32, Design Category=Offline-Designed, Training Protocol=Training-free2025.10 | — | 64.7 | |
| LLaVA-OV-7BFr. (Sampling Rate/Frames)=32, Design Category=Streaming-Designed, Training Protocol=Training-free2025.10 | — | 64.7 | |
| LLaVA-OV-7B + LiveVLMFr. (Sampling Rate/Frames)=0.5/0.2fps, Design Category=Streaming-Designed, Training Protocol=Training-free2025.10 | — | 66.3 | |
| LLaVA-OV-7B + ReKVFr. (Sampling Rate/Frames)=0.5fps, Design Category=Streaming-Designed, Training Protocol=Training-free2025.10 | — | 68.5 | |
| LLaVA-OV-7B + StreamMemFr. (Sampling Rate/Frames)=0.5/0.2fps, Design Category=Streaming-Designed, Training Protocol=Training-free2025.10 | — | 66.9 | |
| LLaVA-Video#S (Model Size)=7B, #F (Number of input frames)=64, Architecture=AR Baseline, reproduced=true2026.01 | — | 70.8 | |
| LLaVA-VideoSize (LLM)=7B (Qwen2), #Visual Tokens=1692026.03 | — | 67.9 | |
| LongVASize (LLM)=7B (Qwen2), #Visual Tokens=1442026.03 | — | 56.3 | |
| LongVAFrames=128, LLM Size=7B2026.06 | — | 56.3 | |
| LongVUSize (LLM)=7B (Qwen2), #Visual Tokens=642026.03 | — | 65.4 | |
| LongVUFrames=1 FPS, LLM Size=7B2026.06 | — | 65.4 | |
| MA-LMMSize (LLM)=7B (Vicuna-v1.1), #Visual Tokens=322026.03 | — | 36.4 | |
| MiMo-VLFrames=32, LLM Size=7B2026.06 | — | 61.2 | |
| MiMo-VL w/ Q-FoldFrames=32, LLM Size=7B2026.06 | — | 67.9 | |
| MovieChatSize (LLM)=7B (Vicuna-v0), #Visual Tokens=322026.03 | — | 25.8 | |
| MovieChat-7BFr. (Sampling Rate/Frames)=2048, Design Category=Streaming-Designed, Training Protocol=Training-based2025.10 | — | 25.8 | |
| mPLUG-Owl3Frames=128, LLM Size=7B2026.06 | — | 70 | |
| OmniVideo-R12026.02 | — | 74.1 |