Video Understanding on LongVideoBench
73.89AccuracyVSI-DETECTOR
Evaluation Results
| Method | Links | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| VSI-DETECTORTraining Required=Training Free, Video Search=YOLO-World-110M, Text Encoding=ALL-MPNET-BASE-V2, Frame=642025.08 | 73.89 | 31.71 | — | — | — | — | — | — | — | — | — | 36.8 | 26.34 | — | |
| VSLS-DETECTORTraining Required=Training Free, Video Search=YOLO-World-110M, Text Encoding=N/A, Frame=642025.08 | 70.23 | 33.26 | — | — | — | — | — | — | — | — | — | 33.3 | 24.49 | — | |
| QCALLM Size=30B, Frames=64, Backbone=Qwen3-VL-30B-A3B2026.07 | 69.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| T*-DETECTORTraining Required=Training Free, Video Search=YOLO-World-110M, Text Encoding=N/A, Frame=642025.08 | 67.58 | 28.96 | — | — | — | — | — | — | — | — | — | 31.7 | 24.76 | — | |
| Qwen3-VL-30B-A3BLLM Size=30B, Frames=642026.07 | 67.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| QCALLM Size=8B, Frames=64, Backbone=Qwen3-VL-8B2026.07 | 66.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4oEvaluation Protocol=vanilla offline2026.05 | 66.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4oSetting=best-setting numbers from official reports2026.05 | 66.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4oLLM Size=-, Frames=256 / 0.5fps2026.07 | 66.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL3.5-38BLLM Size=38B, Frames=642026.07 | 65.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Kimi-VL-16B-A3BLLM Size=16B, Frames=642026.07 | 64.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Gemini 1.5 ProSetting=best-setting numbers from official reports2026.05 | 64.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CapRL-Video-178KBackbone=Molmo2-8B2026.06 | 64.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Gemini-1.5-proEvaluation Protocol=vanilla offline2026.05 | 64 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Gemini-1.5-ProLLM Size=-, Frames=-2026.07 | 64 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-Video-72BLLM Size=72B, Frames=642026.07 | 63.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4o-Video-178KBackbone=Molmo2-8B2026.06 | 63.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8BLLM Size=8B, Frames=642026.07 | 63.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CapRL-Video-178KBackbone=Molmo2-4B2026.06 | 62.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Tarsier2-7B-Video-178KBackbone=Molmo2-8B2026.06 | 62.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8B + ADPOPost-training Strategy=ADPO2026.05 | 62.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7B#Frames=1fps2025.08 | 61.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| GPT-4o-Video-178KBackbone=Molmo2-4B2026.06 | 61.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL-V3.5-8BEvaluation Protocol=vanilla offline2026.05 | 61 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Tarsier2-7B-Video-178KBackbone=Molmo2-4B2026.06 | 61 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ShareGPT4Video-8B-Video-178KBackbone=Molmo2-8B2026.06 | 61 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-72BLLM Size=72B, Frames=1fps2026.07 | 60.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8B + GRPOPost-training Strategy=GRPO2026.05 | 60.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ParaVT-8BSetting=tool-augmented setting (<think>→<tool_call>→<answer>)2026.05 | 60.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ShareGPT4Video-8B-Video-178KBackbone=Molmo2-4B2026.06 | 60.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8B + DAPOPost-training Strategy=DAPO2026.05 | 60.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL-V2.5-8BEvaluation Protocol=vanilla offline2026.05 | 60 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7BBase Model=Qwen2.5-VL-7B, Input Frames=322026.04 | 59.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Video-LLaMA3-7BEvaluation Protocol=vanilla offline2026.05 | 59.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| mPLUG-Owl3LLM Size=8B, Frames=1282026.07 | 59.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-Video 7BRetention Ratio R=100%, # Newline Tokens M=8322026.05 | 59.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| EvoStreaming-8B-InternVL-V3.5Evaluation Protocol=vanilla offline2026.05 | 59.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-Video-7BFrames per video=64, Retain Tokens=All 64 x 169 (100%)2026.07 | 58.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DivPruneFrames per video=64, Retain Tokens=64 x 64 (↓ 62.1%)2026.07 | 58.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| ApolloLLM Size=7B, Frames=2fps2026.07 | 58.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7BEvaluation Protocol=vanilla offline2026.05 | 58.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| FastVRetention Ratio R=100%/25%, # Newline Tokens M=832/568.72026.05 | 58.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-Video-7BEvaluation Protocol=vanilla offline2026.05 | 58.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| EvoStreaming-8B-Qwen2.5Evaluation Protocol=vanilla offline2026.05 | 58.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| StreamAgent-7B#Frames=1fps2025.08 | 57.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DyCokeRetention Ratio R=32.1%, # Newline Tokens M=2562026.05 | 57.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| EADPFrames per video=64, Retain Tokens=64 x 64 (↓ 62.1%)2026.07 | 57.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| TimeChat-Online-7B#Frames=1fps2025.08 | 57.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| NVILALLM Size=7B, Frames=2562026.07 | 57.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-Video-7BBase Model=LLaVA-Video-7B, Input Frames=642026.04 | 57.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Video-R1-7BSetting=reasoning-enhanced setting (<think>→<answer>)2026.05 | 57.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CDPrunerFrames per video=64, Retain Tokens=64 x 64 (↓ 62.1%)2026.07 | 57.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| TimeChat-Online-7BEvaluation Protocol=vanilla offline2026.05 | 57.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DynaTokRetention Ratio R=24.9%, # Newline Tokens M=8322026.05 | 57 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-OneVision-7BRetention Ratio=100%, Sampled Frames=322026.04 | 56.8 | — | — | — | — | — | — | — | — | 59 | 100 | — | — | — | |
| TangoRetention Ratio=20%, Base Model=LLaVA-OneVision-7B, Sampled Frames=322026.04 | 56.7 | — | — | — | — | — | — | — | — | 58.9 | 99.7 | — | — | — | |
| EADPFrames per video=64, Retain Tokens=64 x 32 (↓ 81.1%)2026.07 | 56.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| HoliTomRetention Ratio=15%, Base Model=LLaVA-OneVision-7B, Sampled Frames=322026.04 | 56.5 | — | — | — | — | — | — | — | — | 57.8 | 97.9 | — | — | — | |
| HoliTomRetention Ratio=20%, Base Model=LLaVA-OneVision-7B, Sampled Frames=322026.04 | 56.5 | — | — | — | — | — | — | — | — | 58.2 | 98.6 | — | — | — | |
| TangoRetention Ratio=15%, Base Model=LLaVA-OneVision-7B, Sampled Frames=322026.04 | 56.4 | — | — | — | — | — | — | — | — | 58.5 | 99.1 | — | — | — | |
| Qwen3-VL-8BPost-training Strategy=Base Model2026.05 | 56.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LLaVA-OneVision-7BEvaluation Protocol=vanilla offline2026.05 | 56.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| TangoRetention Ratio=10%, Base Model=LLaVA-OneVision-7B, Sampled Frames=322026.04 | 56.2 | — | — | — | — | — | — | — | — | 58.4 | 98.9 | — | — | — | |
| FastVIDRetention Ratio=20%, Base Model=LLaVA-OneVision-7B, Sampled Frames=322026.04 | 56.2 | — | — | — | — | — | — | — | — | 58.5 | 99 | — | — | — | |
| Time-R1-7BSetting=reasoning-enhanced setting (<think>→<answer>)2026.05 | 56 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Video-Thinker-7BSetting=reasoning-enhanced setting (<think>→<answer>)2026.05 | 56 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| HiPruneFrames per video=64, Retain Tokens=64 x 64 (↓ 62.1%)2026.07 | 56 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| FastVIDRetention Ratio=15%, Base Model=LLaVA-OneVision-7B, Sampled Frames=322026.04 | 55.9 | — | — | — | — | — | — | — | — | 58 | 98.3 | — | — | — | |
| FastVIDBase Model=Qwen2.5-VL-7B, Retention Ratio=20%, Input Frames=322026.04 | 55.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| TangoBase Model=Qwen2.5-VL-7B, Retention Ratio=20%, Input Frames=322026.04 | 55.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| FastVIDRetention Ratio=10%, Base Model=LLaVA-OneVision-7B, Sampled Frames=322026.04 | 55.7 | — | — | — | — | — | — | — | — | 56.9 | 96.4 | — | — | — | |
| VisionZipBase Model=Qwen2.5-VL-7B, Retention Ratio=20%, Input Frames=322026.04 | 55.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DivPruneFrames per video=64, Retain Tokens=64 x 32 (↓ 81.1%)2026.07 | 55.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| TangoBase Model=LLaVA-Video-7B, Retention Ratio=20%, Input Frames=642026.04 | 55.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Qwen2-VL-7BEvaluation Protocol=vanilla offline2026.05 | 55.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| HoliTomRetention Ratio=10%, Base Model=LLaVA-OneVision-7B, Sampled Frames=322026.04 | 55.4 | — | — | — | — | — | — | — | — | 57.1 | 96.7 | — | — | — | |
| Tango†Base Model=LLaVA-Video-7B, Retention Ratio=20%, Intra-LLM=true, Input Frames=642026.04 | 55.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| FastVRetention Ratio R=100%/10%, # Newline Tokens M=832/311.82026.05 | 55.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CDPrunerFrames per video=64, Retain Tokens=64 x 32 (↓ 81.1%)2026.07 | 55.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DynaTokRetention Ratio R=9.5%, # Newline Tokens M=8322026.05 | 55.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VideoRFT-7BSetting=reasoning-enhanced setting (<think>→<answer>)2026.05 | 55.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| HiPruneFrames per video=64, Retain Tokens=64 x 32 (↓ 81.1%)2026.07 | 54.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VisionZipRetention Ratio=20%, Base Model=LLaVA-OneVision-7B, Sampled Frames=322026.04 | 54.8 | — | — | — | — | — | — | — | — | 56.7 | 96 | — | — | — | |
| FastVRetention Ratio=20%, Base Model=LLaVA-OneVision-7B, Sampled Frames=322026.04 | 54.7 | — | — | — | — | — | — | — | — | 54.8 | 92.9 | — | — | — | |
| Tango†Base Model=LLaVA-Video-7B, Retention Ratio=10%, Intra-LLM=true, Input Frames=642026.04 | 54.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LongVT-RFT-7BSetting=tool-augmented setting (<think>→<tool_call>→<answer>)2026.05 | 54.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VisionZipBase Model=LLaVA-Video-7B, Retention Ratio=20%, Input Frames=642026.04 | 54.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| InternVL-V2-8BEvaluation Protocol=vanilla offline2026.05 | 54.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| FastVIDBase Model=LLaVA-Video-7B, Retention Ratio=20%, Input Frames=642026.04 | 54.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Conan-7BSetting=tool-augmented setting (<think>→<tool_call>→<answer>)2026.05 | 54.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| EADPFrames per video=64, Retain Tokens=64 x 16 (↓ 90.5%)2026.07 | 54.5 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| HoliTomBase Model=LLaVA-Video-7B, Retention Ratio=10%, Input Frames=642026.04 | 54.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| TangoBase Model=LLaVA-Video-7B, Retention Ratio=10%, Input Frames=642026.04 | 54.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| HoliTomBase Model=LLaVA-Video-7B, Retention Ratio=20%, Input Frames=642026.04 | 54.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Kangaroo-7BEvaluation Protocol=vanilla offline2026.05 | 54.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VisionZipRetention Ratio R=24.9%, # Newline Tokens M=642026.05 | 54.1 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| CDPrunerFrames per video=64, Retain Tokens=64 x 16 (↓ 90.5%)2026.07 | 53.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| VidCom^2Retention Ratio=20%, Base Model=LLaVA-OneVision-7B, Sampled Frames=322026.04 | 53.7 | — | — | — | — | — | — | — | — | 56.8 | 96.3 | — | — | — | |
| ReWatch-R1-7BSetting=reasoning-enhanced setting (<think>→<answer>)2026.05 | 53.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| EvoStreaming-8B-Qwen2Evaluation Protocol=vanilla offline2026.05 | 53.4 | — | — | — | — | — | — | — | — | — | — | — | — | — |