Video Question Answering on LongVideoBench (val)
83AccuracyHAVEN
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| HAVENfps=0.672026.01 | 83 | — | — | — | — | — | |
| HAVENfps=0.672026.01 | 81 | — | — | — | — | — | |
| Seed1.5-VL-Thinking-200Bfps=0.672026.01 | 74 | — | — | — | — | — | |
| GPT-5.4 (ref.)Type=Proprietary / large-scale reference2026.05 | 72.5 | — | — | — | — | — | |
| DVDfps=0.672026.01 | 71.6 | — | — | — | — | — | |
| Gemini-3-Pro (ref.)Type=Proprietary / large-scale reference2026.05 | 71 | — | — | — | — | — | |
| Qwen3.5-397B-A17 (ref.)Type=Proprietary / large-scale reference2026.05 | 70.5 | — | — | — | — | — | |
| DVDfps=0.672026.01 | 68.6 | — | — | — | — | — | |
| Qwen3.5-122B-A10B (ref.)Type=Proprietary / large-scale reference2026.05 | 68 | — | — | — | — | — | |
| OpenAI o3fps=0.672026.01 | 67.5 | — | — | — | — | — | |
| AdaReTakefps=0.672026.01 | 67 | — | — | — | — | — | |
| GPT-4ofps=0.672026.01 | 66.7 | — | — | — | — | — | |
| GPT-4oParams=-, Frames=384/256/0.5fps2026.03 | 66.7 | — | — | — | — | — | |
| GPT-4o#params=−, #frames=−2026.04 | 66.7 | 60.02 | — | — | — | — | |
| GPT-4oType=Proprietary, Model=–, Frames=2562026.05 | 66.7 | — | — | — | — | — | |
| T*Type=Training-free, Model=LLaVA-OneVision-72B, Frames=322026.05 | 65.4 | — | — | — | — | — | |
| Qwen-2.5-VL#params=7B, #frames=<=7682026.04 | 65.1 | 48.82 | — | — | — | — | |
| MIRAType=Training-free, Model=LLaVA-Video-7B, Frames=642026.05 | 64.5 | — | — | — | — | — | |
| InternVL3.5 + ReFoCUSLLM Size=8B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 64.1 | — | — | — | — | — | |
| Gemini-1.5-ProType=Proprietary, Model=–, Frames=2562026.05 | 64 | — | — | — | — | — | |
| Qwen3-VL + ReFoCUSLLM Size=8B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 63.3 | — | — | — | — | — | |
| A.I.R. + InternVL-3LLM Size=8B, #Frames=≤ 322025.10 | 62.8 | — | — | — | — | — | |
| InternVL3.5 + ReFoCUSLLM Size=4B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 62.6 | — | — | — | — | — | |
| HY-HimmelBackbone=InternVL3-8B2026.05 | 62.5 | — | — | — | — | — | |
| LLaVA-Video + EFSParams=7B, Frames=642026.03 | 62.1 | — | — | — | — | — | |
| InternVL3 + ReFoCUSLLM Size=8B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 62 | — | — | — | — | — | |
| Qwen3-VL + ReFoCUSLLM Size=4B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 61.9 | — | — | — | — | — | |
| Gemini 1.5 Flash#params=−, #frames=−2026.04 | 61.6 | — | — | — | — | — | |
| Gemini-1.5-FlashType=Proprietary, Model=–, Frames=2562026.05 | 61.6 | — | — | — | — | — | |
| TTA-Vidbase_model=InternVL-3, #params=8B, #frames=322026.04 | 61.48 | 55.13 | — | — | — | — | |
| A.I.R. + QwenVL-2.5LLM Size=7B, #Frames=≤ 322025.10 | 61.4 | — | — | — | — | — | |
| GPT-4VType=Proprietary, Model=–, Frames=2562026.05 | 61.3 | — | — | — | — | — | |
| HY-HimmelBackbone=Qwen2.5-VL-7B2026.05 | 61 | — | — | — | — | — | |
| LLaVA-OneVision + ReFoCUSLLM Size=7B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 61 | — | — | — | — | — | |
| GPT-4ofps=0.672026.01 | 60.9 | — | — | — | — | — | |
| InternVL-3 + MDP3LLM Size=8B, #Frames=32, Source=Reproduced2025.10 | 60.9 | — | — | — | — | — | |
| Qwen2.5-VL-72Bfps=0.672026.01 | 60.7 | — | — | — | — | — | |
| A.I.R. + LLaVA-OneVisionLLM Size=7B, #Frames=≤ 322025.10 | 60.7 | — | — | — | — | — | |
| OpenAI o3fps=0.672026.01 | 60.6 | — | — | — | — | — | |
| Qwen2.5-VL + EFSParams=7B, Frames=162026.03 | 60.5 | — | — | — | — | — | |
| LLaVA-OneVision + EFSParams=7B, Frames=82026.03 | 60.3 | — | — | — | — | — | |
| CATS (ours)†Type=Training-free, Model=LLaVA-Video-7B, Frames=322026.05 | 60.21 | — | — | — | — | — | |
| TPOType=Training-based, Model=LLaVA-Video-7B, Frames=642026.05 | 60.1 | — | — | — | — | — | |
| QwenVL-2.5 + MDP3LLM Size=7B, #Frames=32, Source=Reproduced2025.10 | 60 | — | — | — | — | — | |
| HY-HimmelBackbone=LLaVA-Video-7B2026.05 | 59.9 | — | — | — | — | — | |
| AKS†Type=Training-free, Model=LLaVA-Video-7B, Frames=322026.05 | 59.76 | — | — | — | — | — | |
| InternVL3.5LLM Size=8B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 59.7 | — | — | — | — | — | |
| LLaVA-OneVision + BOLTLLM Size=7B, #Frames=32, Source=Reported2025.10 | 59.6 | — | — | — | — | — | |
| DToMAType=Training-free, Model=LLaVA-Video-7B, Frames=642026.05 | 59.6 | — | — | — | — | — | |
| BIMBAType=Training-based, Model=BIMBA-7B, Frames=1282026.05 | 59.5 | — | — | — | — | — | |
| LLaVA-OneVision + AKSLLM Size=7B, #Frames=32, Source=Reported2025.10 | 59.3 | — | — | — | — | — | |
| Video-R2#params=7B, #frames=1282026.04 | 59.2 | 53.92 | — | — | — | — | |
| InternVL-3#params=8B, #frames=322026.04 | 59.08 | 51.37 | — | — | — | — | |
| LLaVA-OneVision + MDP3LLM Size=7B, #Frames=32, Source=Reproduced2025.10 | 59 | — | — | — | — | — | |
| InternVL3-8B2026.05 | 58.9 | — | — | — | — | — | |
| LLaVA-VideoParams=7B, Frames=642026.03 | 58.8 | — | — | — | — | — | |
| InternVL-3 (reported)LLM Size=8B, #Frames=max: 64, Source=Reported2025.10 | 58.8 | — | — | — | — | — | |
| ApolloType=Training-based, Model=Apollo-7B, Frames=2FPS2026.05 | 58.5 | — | — | — | — | — | |
| InternVL-3 (reproduced)LLM Size=8B, #Frames=32, Source=Reproduced2025.10 | 58.3 | — | — | — | — | — | |
| LLaVA-VideoType=Foundational, Model=LLaVA-Video-7B, Frames=642026.05 | 58.2 | — | — | — | — | — | |
| QwenVL-2.5 (reproduced)LLM Size=7B, #Frames=32, Source=Reproduced2025.10 | 58.1 | — | — | — | — | — | |
| TTA-Vidbase_model=Qwen2.5-VL, #params=7B, #frames=322026.04 | 57.81 | 51.84 | — | — | — | — | |
| InternVL3LLM Size=8B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 57.8 | — | — | — | — | — | |
| NVILAParams=8B, Frames=2562026.03 | 57.7 | — | — | — | — | — | |
| InternVL3.5LLM Size=4B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 57.7 | — | — | — | — | — | |
| Video-R1#params=7B, #frames=1282026.04 | 57.6 | 51.94 | — | — | — | — | |
| Qwen2.5-VL-7B2026.05 | 57.4 | — | — | — | — | — | |
| Qwen3-VLLLM Size=4B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 57.4 | — | — | — | — | — | |
| LongVILAType=Training-based, Model=LongVILA-7B, Frames=2562026.05 | 57.1 | — | — | — | — | — | |
| TTRV*base_model=Qwen2.5-VL, #params=7B, #frames=322026.04 | 57.07 | 50.46 | — | — | — | — | |
| Qwen2.5-VLParams=7B, Frames=162026.03 | 57 | — | — | — | — | — | |
| Video-RFT#params=7B, #frames=1282026.04 | 57 | 52.32 | — | — | — | — | |
| VideoLLaMA 3 + ReFoCUSLLM Size=7B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 57 | — | — | — | — | — | |
| LLaVA-OneVision (reproduced)LLM Size=7B, #Frames=32, Source=Reproduced2025.10 | 56.6 | — | — | — | — | — | |
| Video-RTS#params=7B, #frames=1282026.04 | 56.6 | — | — | — | — | — | |
| Qwen3-VLLLM Size=8B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 56.6 | — | — | — | — | — | |
| LLaVA-OneVision (reported)LLM Size=7B, #Frames=32*, Source=Reported2025.10 | 56.4 | — | — | — | — | — | |
| LLaVA-OneVision#params=7B, #frames=642026.04 | 56.3 | 43.27 | — | — | — | — | |
| LLaVA-Video-7B2026.05 | 56.2 | — | — | — | — | — | |
| LiveVLMFrames=0.5/0.2 fps2025.05 | 56.1 | — | — | — | — | — | |
| LLaVA-OneVision-7B + ReKVFrames=0.5/0.2 fps2025.05 | 55.8 | — | — | — | — | — | |
| LLaVA-OneVision-7BFrames=322025.05 | 55.6 | — | — | — | — | — | |
| LLaVA-OneVisionLLM Size=7B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 55 | — | — | — | — | — | |
| MiniCPM-V 2.6Type=Foundational, Model=MiniCPM-V 2.6-8B, Frames=642026.05 | 54.9 | — | — | — | — | — | |
| InternVL3 + ReFoCUSLLM Size=2B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 54.9 | — | — | — | — | — | |
| VideoLLaMA 3LLM Size=7B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 54.8 | — | — | — | — | — | |
| LLaVA-OneVision-7B + StreamMemFrames=0.5/0.2 fps2025.05 | 54.4 | — | — | — | — | — | |
| VideoChat-R1#params=7B, #frames=1282026.04 | 54.3 | 52.34 | — | — | — | — | |
| Kangaroo-8BFrames=642025.05 | 54.2 | — | — | — | — | — | |
| LLaVA-OneVisionParams=7B, Frames=82026.03 | 54.1 | — | — | — | — | — | |
| VideoLLaMA 3 + ReFoCUSLLM Size=2B, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 53.7 | — | — | — | — | — | |
| VideoChat-R1.5#params=7B, #frames=1282026.04 | 53.6 | 52.24 | — | — | — | — | |
| PLLaVAType=Foundational, Model=PLLaVA-34B, Frames=322026.05 | 53.2 | — | — | — | — | — | |
| A.I.R. + VILA-1.5LLM Size=8B, #Frames=82025.10 | 52.9 | — | — | — | — | — | |
| GPT-4o + ReFoCUSLLM Size=Closed, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 52.9 | — | — | — | — | — | |
| VILA-1.5 + MDP3LLM Size=8B, #Frames=8, Source=Reproduced2025.10 | 52.3 | — | — | — | — | — | |
| VILA-1.5 + Q-FrameLLM Size=8B, #Frames=8, Source=Reported2025.10 | 51.6 | — | — | — | — | — | |
| Gemini 2.5 Flash + ReFoCUSLLM Size=Closed, Sampling Strategy=ReFoCUS, Frame Budget=32-frame2025.06 | 50.9 | — | — | — | — | — | |
| InternVL3LLM Size=2B, Sampling Strategy=Uniform sampling, Frame Budget=32-frame2025.06 | 50.9 | — | — | — | — | — | |
| Video-XLParams=7B, Frames=128/2562026.03 | 50.7 | — | — | — | — | — |