Video Question-Answering on LVBench
84.1AccuracyHAVEN
Evaluation Results
| Method | Links | |
|---|---|---|
| HAVENfps=22026.01 | 84.1 | |
| DVD w. subtitlefps=2, subtitles=true2026.01 | 76 | |
| DVDfps=22026.01 | 74.2 | |
| Seed1.5-VL-Thinking-200Bfps=22026.01 | 64.6 | |
| VideoLucyfps=22026.01 | 58.8 | |
| OpenAI o3fps=22026.01 | 57.1 | |
| AdaReTakefps=22026.01 | 53.3 | |
| VideoRAGfps=22026.01 | 49.2 | |
| GPT-4ofps=22026.01 | 48.9 | |
| POINTS-LongNum Frame=504+8, Token/Frame=16, Total Num of Token=93442026.04 | 48.6 | |
| Qwen2.5-VL-72Bfps=22026.01 | 47.7 | |
| VideoChat-FlashSize=7B, #Tokens=162025.01 | 47.2 | |
| InternVideo2.5 (InternVL2.5+LRC)Size=7B, #Tokens=162025.01 | 46.4 | |
| POINTS-LongNum Frame=248+8, Token/Frame=16, Total Num of Token=52482026.04 | 46.3 | |
| POINTS-LongNum Frame=504+8, Token/Frame=8, Total Num of Token=52482026.04 | 46 | |
| VanillaBackbone=Qwen3-VL-8B, Retention Ratio R=100%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 45.7 | |
| POINTS-LongNum Frame=248+8, Token/Frame=8, Total Num of Token=32002026.04 | 44.2 | |
| ResAdaptBackbone=Qwen3-VL-8B, Retention Ratio R=22.9%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 43.3 | |
| ToMeBackbone=Qwen3-VL-8B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 43.1 | |
| Weaver#Frames=128+1fps2026.02 | 43 | |
| ResAdaptBackbone=Qwen3-VL-8B, Retention Ratio R=11.1%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 42.6 | |
| Random DropBackbone=Qwen3-VL-8B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 42.4 | |
| ToMeBackbone=Qwen2.5-VL-7B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 42.1 | |
| ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=22.9%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 42.1 | |
| VanillaBackbone=Qwen2.5-VL-7B, Retention Ratio R=100%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 42 | |
| Flash-VstreamTotal Num of Token=115202026.04 | 42 | |
| VideoAuto-R1 + ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=23.8%, Reasoning (CoT)=true, Frames=128 Frames2026.03 | 41.8 | |
| FixedScaleBackbone=Qwen3-VL-8B, Retention Ratio R=12.3%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 41.7 | |
| POINTS1.5-8B-onlineNum Frame=64, Token/Frame=324, Total Num of Token=207362026.04 | 41.7 | |
| VideoAuto-R1Backbone=Qwen2.5-VL-7B, Retention Ratio R=100%, Reasoning (CoT)=true, Frames=32 Frames2026.03 | 41.5 | |
| VisionZipBackbone=Qwen2.5-VL-7B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 41.5 | |
| QwenVL2Size=72B2025.01 | 41.3 | |
| Random DropBackbone=Qwen2.5-VL-7B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 41.3 | |
| VideoAuto-R1Backbone=Qwen2.5-VL-7B, Retention Ratio R=100%, Reasoning (CoT)=true, Frames=128 Frames2026.03 | 41.2 | |
| VisionZipBackbone=Qwen3-VL-8B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 41.2 | |
| VideoRFT#Frames=322026.02 | 41.1 | |
| VideoMindSize=7B2025.03 | 40.8 | |
| Qwen2.5-VL#Frames=1282026.02 | 40.6 | |
| Video-R1#Frames=642026.02 | 40.5 | |
| ToMeBackbone=Qwen3-VL-8B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 40.5 | |
| FixedScaleBackbone=Qwen2.5-VL-7B, Retention Ratio R=12.3%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 40.3 | |
| VanillaBackbone=Qwen3-VL-8B, Retention Ratio R=100%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 40.2 | |
| Weaver-SFT#Frames=128+1fps2026.02 | 40.1 | |
| ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=11.1%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 39.9 | |
| Qwen2-VL-onlineTotal Num of Token=115202026.04 | 39.8 | |
| ToMeBackbone=Qwen2.5-VL-7B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 39.5 | |
| VideoAuto-R1 + ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=11.4%, Reasoning (CoT)=true, Frames=128 Frames2026.03 | 39.5 | |
| VisionZipBackbone=Qwen3-VL-8B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 39.4 | |
| PixelReasonser#Frames=162026.02 | 39.2 | |
| Random DropBackbone=Qwen3-VL-8B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 38.8 | |
| VideoAuto-R1 + ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=23.8%, Reasoning (CoT)=true, Frames=32 Frames2026.03 | 38.7 | |
| VanillaBackbone=Qwen2.5-VL-7B, Retention Ratio R=100%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 38.6 | |
| FlashVidBackbone=Qwen3-VL-8B, Retention Ratio R=30.0%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 38.5 | |
| ResAdaptBackbone=Qwen3-VL-8B, Retention Ratio R=23.8%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 38.5 | |
| InternVL2.5Size=7B, #Tokens=2562025.01 | 38.4 | |
| InternVL2.5#Frames=16-642026.02 | 38.4 | |
| VisionZipBackbone=Qwen2.5-VL-7B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 38.2 | |
| Random DropBackbone=Qwen2.5-VL-7B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=128 Frames2026.03 | 38 | |
| ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=23.8%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 37.9 | |
| Random DropBackbone=Qwen3-VL-8B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 37.8 | |
| ToMeBackbone=Qwen2.5-VL-7B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 37.7 | |
| FixedScaleBackbone=Qwen2.5-VL-7B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 37.3 | |
| VisionZipBackbone=Qwen3-VL-8B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 37.3 | |
| ResAdaptBackbone=Qwen3-VL-8B, Retention Ratio R=11.4%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 37.3 | |
| FlashVidBackbone=Qwen3-VL-8B, Retention Ratio R=12.2%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 37.1 | |
| FixedScaleBackbone=Qwen3-VL-8B, Retention Ratio R=12.3%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 37.1 | |
| FlashVidBackbone=Qwen2.5-VL-7B, Retention Ratio R=29.3%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 36.9 | |
| Random DropBackbone=Qwen2.5-VL-7B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 36.7 | |
| VideoAuto-R1 + ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=11.4%, Reasoning (CoT)=true, Frames=32 Frames2026.03 | 36.7 | |
| VisionZipBackbone=Qwen2.5-VL-7B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 36.5 | |
| FlashVidBackbone=Qwen2.5-VL-7B, Retention Ratio R=10.4%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 36.5 | |
| ToMeBackbone=Qwen3-VL-8B, Retention Ratio R=25.0%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 36.4 | |
| ResAdaptBackbone=Qwen2.5-VL-7B, Retention Ratio R=11.4%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 35.9 | |
| ToMeBackbone=Qwen3-VL-8B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 35.8 | |
| Random DropBackbone=Qwen3-VL-8B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 35.5 | |
| VideoMindSize=2B2025.03 | 35.4 | |
| FixedScaleBackbone=Qwen2.5-VL-7B, Retention Ratio R=12.3%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 35.4 | |
| VisionZipBackbone=Qwen3-VL-8B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 35.4 | |
| VisionZipBackbone=Qwen2.5-VL-7B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 35.3 | |
| Random DropBackbone=Qwen2.5-VL-7B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 35.2 | |
| ToMeBackbone=Qwen2.5-VL-7B, Retention Ratio R=10.0%, Reasoning (CoT)=false, Frames=32 Frames2026.03 | 33.6 | |
| Gemini-1.5-Pro2025.01 | 33.1 | |
| Gemini-1.5-Pro2025.03 | 33.1 | |
| Gemini-1.5-Pro#Frames=1fps2026.02 | 33.1 | |
| GPT4-o2025.01 | 30.8 | |
| GPT-4o2025.03 | 30.8 | |
| GPT4-o#Frames=1fps2026.02 | 30.8 | |
| VideoAgentfps=22026.01 | 29.3 | |
| VideoTreefps=22026.01 | 28.8 | |
| LLaVA-OneVisionSize=72B, #Tokens=1962025.01 | 26.9 | |
| PLLAVASize=34B2025.03 | 26.1 | |
| VimRAGBackbone=Qwen3-VL-8B-Instruct2026.02 | 24.5 | |
| LLaMA-VIDSize=7B, #Tokens=22025.01 | 23.9 | |
| VideoRAGBackbone=Qwen3-VL-8B-Instruct2026.02 | 23.8 | |
| VimRAGBackbone=Qwen3-VL-4B-Instruct2026.02 | 22.8 | |
| MovieChatSize=7B2025.03 | 22.5 | |
| Mem1Backbone=Qwen3-VL-8B-Instruct2026.02 | 22.4 | |
| TimeChatSize=7B2025.03 | 22.3 | |
| MemAgentBackbone=Qwen3-VL-8B-Instruct2026.02 | 22.2 | |
| VideoRAGBackbone=Qwen3-VL-4B-Instruct2026.02 | 19.4 |