Long Video Understanding on LongVideoBench
76.8ScoreGemini 2.5 Pro
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Gemini 2.5 ProModel Type=Proprietary2026.06 | 76.8 | — | |
| Gemini 3 ProModel Type=Proprietary2026.06 | 75.9 | — | |
| Keye-VL-2.0-30B-A3BModel Parameters=30B-A3B2026.06 | 74.1 | — | |
| Gemini 2.5 FlashModel Type=Proprietary2026.06 | 73.1 | — | |
| GPT-5Model Type=Proprietary2026.06 | 72.6 | — | |
| Qwen2.5-VL w/ AdaptTokenLLM Size=72B2026.03 | 70.5 | — | |
| Qwen3-VL ThinkingModel Parameters=235B-A22B, Thinking Mode=true2026.06 | 70.5 | — | |
| ByteVideoLLMLLM Size=14B2026.03 | 70.1 | — | |
| GPT-5 miniModel Type=Proprietary2026.06 | 69.7 | — | |
| Molmo2-8BModel Type=Open-Weight, Parameters=8B2026.06 | 67.5 | — | |
| InternVL3.5Model Parameters=241B-A28B2026.06 | 67.1 | — | |
| Qwen2.5-VL w/ AdaReTAKELLM Size=72B2026.03 | 67 | — | |
| Qwen3-VL w/ AdaptTokenLLM Size=8B2026.03 | 66.9 | — | |
| InternVideo3Model Type=Open-Weight2026.06 | 66.8 | — | |
| GPT-4o2024.12 | 66.7 | — | |
| GPT-4oLLM Size=-2026.03 | 66.7 | — | |
| GPT-4oEvaluation Protocol=Zero-shot2025.07 | 66.7 | — | |
| Qwen2.5-VL w/ FlexSelectLLM Size=72B2026.03 | 66.4 | — | |
| Eagle2.5-8BModel Type=Open-Weight, Parameters=8B2026.06 | 66.4 | — | |
| Qwen2.5-VLLLM Size=72B2026.03 | 66.2 | — | |
| Keye-VL-1.5-8BModel Type=Open-Weight, Parameters=8B2026.06 | 66 | — | |
| Qwen3-VLLLM Size=8B2026.03 | 65.9 | — | |
| GLM-4.1V-9BModel Type=Open-Weight, Parameters=9B2026.06 | 65.7 | — | |
| Qwen2.5-VL w/ AdaptTokenLLM Size=7B2026.03 | 65.2 | — | |
| Qwen2.5-VL w/ AdaptToken-LiteLLM Size=7B2026.03 | 65.1 | — | |
| Claude Sonnet 4.5Model Type=Proprietary2026.06 | 65.1 | — | |
| VideoChat-Flash-7BModel Type=Open-Weight, Parameters=7B2026.06 | 64.7 | — | |
| Video-CCAMSize=9B2024.12 | 64.6 | — | |
| DIG#Frames=768, Backbone=Qwen3-VL-8B2025.12 | 64.6 | — | |
| Gemini-1.5-Pro2024.12 | 64 | — | |
| Gemini-1.5-ProLLM Size=-2026.03 | 64 | — | |
| Gemini-1.5-ProEvaluation Protocol=Zero-shot2025.07 | 64 | — | |
| Qwen2.5-VL w/ SeViCESLLM Size=7B2026.03 | 63.9 | — | |
| MiniCPM-V-4.5-8BModel Type=Open-Weight, Parameters=8B2026.06 | 63.9 | — | |
| DIG#Frames=512, Backbone=Qwen3-VL-8B2025.12 | 63.8 | — | |
| InternVL2.5 w/ AdaptToken-LiteLLM Size=8B2026.03 | 63.8 | — | |
| InternVL2.5 w/ AdaptTokenLLM Size=8B2026.03 | 63.7 | — | |
| InternVL2.5 w/ ZoomVLLM Size=8B2026.03 | 63.3 | — | |
| Video-CCAMSize=4B2024.12 | 62.8 | — | |
| Qwen3-VL-8B-InstructMax Input Frames=642026.03 | 62.8 | — | |
| Qwen2.5-VL w/ AdaReTAKELLM Size=7B2026.03 | 62.6 | — | |
| Qwen3-VL-8BRetention Ratio γ=100%2026.04 | 62.5 | — | |
| Qwen3-VL-32B-InstructInput Frames=322026.03 | 62.4 | — | |
| Qwen2.5-VL w/ FlexSelectLLM Size=7B2026.03 | 62.4 | — | |
| Qwen3-VL-8BModel Type=Open-Weight, Parameters=8B2026.06 | 62.4 | — | |
| InternVL3.5-8BModel Type=Open-Weight, Parameters=8B2026.06 | 62.1 | — | |
| InternVL2.5 w/ SeViCESLLM Size=8B2026.03 | 61.7 | — | |
| V-CASTMax Input Frames=64, Retention Ratio=35%2026.03 | 61.6 | — | |
| Gemini-1.5-FlashEvaluation Protocol=Zero-shot2025.07 | 61.6 | — | |
| Qwen3.5Model Parameters=35B-A3B2026.06 | 61.6 | — | |
| V-CASTBase Model=Qwen3-VL-32B-Instruct, Input Frames=32, Retention Ratio=25%2026.03 | 61.5 | — | |
| KiTokeRetention Ratio γ=25%2026.04 | 61.5 | — | |
| GPT-4V2024.12 | 61.3 | — | |
| DIG#Frames=256, Backbone=Qwen3-VL-8B2025.12 | 61.2 | — | |
| ViLAMPLLM Size=7B2026.03 | 61.2 | — | |
| VidCom2Max Input Frames=64, Retention Ratio=35%2026.03 | 61 | — | |
| Qwen2.5-VL w/ ZoomVLLM Size=7B2026.03 | 61 | — | |
| UNI#Frames=768, Backbone=Qwen3-VL-8B2025.12 | 60.9 | — | |
| FastVIDBase Model=Qwen3-VL-32B-Instruct, Input Frames=32, Retention Ratio=25%2026.03 | 60.9 | — | |
| ReMoRaLLM backbone=Qwen2-7B2026.02 | 60.8 | — | |
| FastVIDMax Input Frames=64, Retention Ratio=35%2026.03 | 60.7 | — | |
| InternVL2.5 w/ TriumphLLM Size=8B2026.03 | 60.7 | — | |
| InternVideo2.5-7BModel Type=Open-Weight, Parameters=7B2026.06 | 60.6 | — | |
| DIG#Frames=192, Backbone=Qwen3-VL-8B2025.12 | 60.4 | — | |
| VidCom2Base Model=Qwen3-VL-32B-Instruct, Input Frames=32, Retention Ratio=25%2026.03 | 60.4 | — | |
| HoliTomMax Input Frames=64, Retention Ratio=35%2026.03 | 60.4 | — | |
| VisionZipRetention Ratio γ=25%2026.04 | 60.4 | — | |
| Qwen3-VL-8B-InstructMax Input Frames=322026.03 | 60.3 | — | |
| VidCom2Retention Ratio γ=25%2026.04 | 60.3 | — | |
| UNI#Frames=512, Backbone=Qwen3-VL-8B2025.12 | 60.2 | — | |
| HoliTomBase Model=Qwen3-VL-32B-Instruct, Input Frames=32, Retention Ratio=25%2026.03 | 60.2 | — | |
| TPOLLM Size=7B2026.03 | 60.1 | — | |
| InternVL2.5 w/ FlexSelectLLM Size=8B2026.03 | 60.1 | — | |
| Qwen2.5-VL w/ TimeSearch-RLLM Size=7B2026.03 | 60.1 | — | |
| VisionZipBase Model=Qwen3-VL-32B-Instruct, Input Frames=32, Retention Ratio=25%2026.03 | 60 | — | |
| VideoLLaMA3LLM Size=7B2026.03 | 59.8 | — | |
| VideoLLaMA3LLM=Qwen2.5-7B, Frames=1fps, Evaluation Protocol=Zero-shot2025.07 | 59.8 | — | |
| FastVIDRetention Ratio γ=25%2026.04 | 59.7 | — | |
| BIMBALLM backbone=Qwen2-7B2026.02 | 59.5 | — | |
| Qwen2.5-VLLLM backbone=Qwen2.5-7B2026.02 | 59.5 | — | |
| VanillaBase Model=LLaVA-Video, Retention Ratio R=100%2026.02 | 59.5 | — | |
| VanillaRetention Ratio R=100%2026.02 | 59.5 | — | |
| InternVL2.5LLM Size=8B2026.03 | 59.5 | — | |
| Qwen2.5-VLLLM Size=7B2026.03 | 59.5 | — | |
| PVC InternVL2Size=8B, # token/frame=642024.12 | 59.2 | — | |
| FlashVIDMax Input Frames=32, Retention Ratio=35%2026.03 | 59.2 | — | |
| PruneVIDRetention Ratio γ=25%2026.04 | 59.2 | — | |
| FlashVIDRetention Ratio R=25%2026.02 | 59.1 | — | |
| EchoPruneNumber of Frames (F)=320, Compression Ratio=20x, Token Budget=5% (ultra)2026.05 | 59 | — | |
| LLaVA-Video-7BPrefilling FLOPs (T)=80.2, FLOPs Ratio=100%, Before LLM Retained Ratio=100%2026.03 | 58.9 | — | |
| VanillaRetention Ratio R=100%2026.02 | 58.9 | — | |
| LLaVA-Video-7BBackbone=LLaVA-Video-7B, FLOPs (T)=80.9, FLOPs Ratio=100%, Retention Ratio=100%2026.03 | 58.9 | — | |
| VisionZipMax Input Frames=64, Retention Ratio=35%2026.03 | 58.9 | — | |
| FlashVIDBase Model=LLaVA-Video, Retention Ratio R=20%2026.02 | 58.7 | — | |
| FastVIDMax Input Frames=32, Retention Ratio=35%2026.03 | 58.7 | — | |
| VidCom2Max Input Frames=32, Retention Ratio=35%2026.03 | 58.6 | — | |
| KiTokeRetention Ratio γ=10%2026.04 | 58.6 | — | |
| FlashVIDBase Model=LLaVA-OneVision, Retention Ratio R=20%2026.02 | 58.5 | — | |
| V-CASTMax Input Frames=32, Retention Ratio=35%2026.03 | 58.5 | — | |
| DIG#Frames=128, Backbone=Qwen3-VL-8B2025.12 | 58.3 | — |