Long Video Understanding on VideoMME
89.5AccuracySeed2.0 Pro
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Seed2.0 ProSubtitles=true2026.06 | 89.5 | — | |
| Gemini-3-ProSubtitles=true2026.06 | 88.4 | — | |
| Seed1.8Subtitles=true2026.06 | 87.8 | — | |
| Seed2.0 LiteSubtitles=true2026.06 | 87.7 | — | |
| Gemini-3-FlashSubtitles=true2026.06 | 85.2 | — | |
| Gemini-1.5 Prosubtitles=true2024.10 | 81.3 | — | |
| Seed2.0 MiniSubtitles=true2026.06 | 81.2 | — | |
| GPT-4osubtitles=true2024.10 | 77.2 | — | |
| GPT4-O2025.01 | 77.2 | — | |
| Gemini-1.5 Flashsubtitles=true2024.10 | 75 | — | |
| Gemini-1.5-ProSize=-, Evaluation Protocol=Standard2025.12 | 75 | 67.4 | |
| Gemini-1.5-Pro2025.11 | 75 | 67.4 | |
| ARIAsubtitles=true2024.10 | 72.1 | — | |
| GPT-4oSize=-, Evaluation Protocol=Standard2025.12 | 71.9 | 65.3 | |
| GPT-4oPerception budget constraint=false, Perc. budget=-2026.05 | 71.9 | — | |
| GPT4o2025.11 | 71.9 | 65.3 | |
| GPT-4o minisubtitles=true2024.10 | 68.9 | — | |
| VideoLLaMA3-7BPerception budget constraint=false, Perc. budget=-2026.05 | 66.2 | — | |
| StreamReadySize=7B, Backbone=Qwen-2-VL2026.03 | 65.8 | — | |
| REVISOR (Ours)Model Size=7B, Video Tokens=8K2025.11 | 65.7 | 56.2 | |
| VideoChat-FlashModel Size=7B, Video Tokens=8K2025.11 | 65.3 | 55.4 | |
| VideoZoomerSize=7B, Evaluation Protocol=Max 128 frames2025.12 | 65.2 | 55.8 | |
| LongVILA-R1Model Size=7B2025.11 | 65.1 | 55.2 | |
| StreamBridgeSize=7B, Backbone=Qwen-2-VL2026.03 | 64.4 | — | |
| Qwen2.5-VL⋆Model Size=7B, Video Tokens=8K, Status=reproduction2025.11 | 64.3 | 53.4 | |
| NVILAModel Size=8B, Video Tokens=8K2025.11 | 64.2 | 54.8 | |
| HierarQSize=7B2026.03 | 63.7 | — | |
| Hour-LLaVAModel Size=7B2025.11 | 63.6 | 55 | |
| Open-o3-VideoModel Size=7B, Video Tokens=2K2025.11 | 63.6 | 54.9 | |
| Qwen2.5-VLSize=7B, Evaluation Protocol=Max 128 frames2025.12 | 63.5 | 53.9 | |
| Qwen2.5-VL*Model Size=7B, Video Tokens=8K, Training=text-based reflection mechanism2025.11 | 63.4 | 53.2 | |
| GPT-4Vsubtitles=true2024.10 | 63.3 | — | |
| Qwen-2-VLSize=7B2026.03 | 63.3 | — | |
| LLaVA-VideoSize=7B2026.03 | 63.3 | — | |
| InfiniPot-VSize=7B, Backbone=Qwen-2-VL2026.03 | 62.8 | — | |
| TimeChat-OnlineSize=7B2026.03 | 62.5 | — | |
| LongVILA-R1Size=7B, Evaluation Protocol=Standard2025.12 | 62.4 | 53.3 | |
| VL-RethinkerModel Size=7B, Video Tokens=8K2025.11 | 62.1 | 51.9 | |
| StreamForestSize=7B, Backbone=Qwen-2-VL2026.03 | 61.4 | — | |
| Video-R1Model Size=7B, Video Tokens=8K2025.11 | 61.4 | — | |
| Flash-VStreamSize=7B, Backbone=Qwen-2-VL2026.03 | 61.2 | — | |
| Video-R1Size=7B, Evaluation Protocol=Max 128 frames2025.12 | 61.1 | 51.4 | |
| GPT4-V2025.01 | 60.7 | — | |
| LongVUSize=7B, Evaluation Protocol=Standard2025.12 | 60.6 | — | |
| LongVUSize=7B2026.03 | 60.6 | — | |
| LongVUModel Size=7B, Video Tokens=8K2025.11 | 60.6 | 59.5 | |
| MACF (Qwen3-VL-8B)Perception budget constraint=true, Perc. budget=16*224*224, Backbone=Qwen3-VL-8B2026.05 | 60.4 | — | |
| LongVILASize=7B, Evaluation Protocol=Standard2025.12 | 60.1 | — | |
| MACF (LLaVA-OV1.5-8B)Perception budget constraint=true, Perc. budget=16*224*224, Backbone=LLaVA-OneVision1.5-8B2026.05 | 59.9 | — | |
| Qwen2.5-VL-72BPerception budget constraint=true, Perc. budget=16*224*2242026.05 | 59.5 | — | |
| Video-MTRModel Size=7B, Video Tokens=4K2025.11 | 59 | 51 | |
| LLaVA-OneVisionSize=7B, Evaluation Protocol=Standard2025.12 | 58.3 | 46.7 | |
| Qwen3-VL-30BPerception budget constraint=true, Perc. budget=16*224*2242026.05 | 58.3 | — | |
| LLaVA-OneVisionSize=7B2026.03 | 58.2 | — | |
| MACF (Qwen2.5-VL-7B)Perception budget constraint=true, Perc. budget=16*224*224, Backbone=Qwen2.5-VL-7B2026.05 | 58.2 | — | |
| FrameOracleFrame=64 → 15.6, Backbone=LLaVA-OneVision2025.10 | 58.1 | — | |
| VambaModel Size=10B2025.11 | 57.8 | — | |
| DispiderSize=7B, Backbone=Qwen-2-VL2026.03 | 57.2 | — | |
| LLaVA-OneVision1.5-8BPerception budget constraint=true, Perc. budget=16*224*2242026.05 | 56.1 | — | |
| KangarooSize=8B, Evaluation Protocol=Standard2025.12 | 56 | — | |
| Kangaroo-8BPerception budget constraint=false, Perc. budget=-2026.05 | 56 | — | |
| Qwen3-VL-8BPerception budget constraint=true, Perc. budget=16*224*2242026.05 | 55.9 | — | |
| LLaVA-OctopusTraining Data=same as Video-LLaVA [36]2025.01 | 55.7 | — | |
| LatentMASPerception budget constraint=true, Perc. budget=16*224*224, VLM engine=Qwen3-VL-8B2026.05 | 55.7 | — | |
| Keye1.5-VL-8BPerception budget constraint=true, Perc. budget=16*224*2242026.05 | 55.6 | — | |
| Video-XLSize=7B, Evaluation Protocol=Standard2025.12 | 55.5 | — | |
| ViSpeakSize=7B, Backbone=Qwen-2-VL2026.03 | 55 | — | |
| VideoLLaMA2.1-7BPerception budget constraint=false, Perc. budget=-2026.05 | 54.9 | — | |
| LLaVA-Octopus2025.01 | 54.7 | — | |
| VideoChat2-HD2025.01 | 54.6 | — | |
| LongVA2025.01 | 54.3 | — | |
| LongVAModel Size=7B, Video Tokens=224K2025.11 | 54.3 | 47.6 | |
| InternVL2Size=8B2026.03 | 54 | — | |
| Qwen2.5-VL-7BPerception budget constraint=true, Perc. budget=16*224*2242026.05 | 53.7 | — | |
| VideoChat-onlineSize=4B2026.03 | 52.8 | — | |
| LongVASize=7B, Evaluation Protocol=Standard2025.12 | 52.6 | — | |
| LongVASize=7B2026.03 | 52.6 | — | |
| LLaVA-Next-Video-34BPerception budget constraint=false, Perc. budget=-2026.05 | 52 | — | |
| Llama3.2-11Bsubtitles=true2024.10 | 50.2 | — | |
| InternVL2.5-VL-8BPerception budget constraint=true, Perc. budget=16*224*2242026.05 | 47.7 | — | |
| Pixtral-12Bsubtitles=true2024.10 | 47.5 | — | |
| MapReducePerception budget constraint=true, Perc. budget=16*224*224, VLM engine=Qwen3-VL-8B2026.05 | 46.7 | — | |
| LLaVA-OneVisionModel Size=7B, Video Tokens=6K2025.11 | 46.7 | — | |
| VideoLLaMA22025.01 | 46.6 | — | |
| LLaVA-NeXT-Video2025.01 | 46.5 | — | |
| Chat-UniVi2025.01 | 45.9 | — | |
| MovieChatFrame=512 → 64, Backbone=LLaVA-OneVision2025.10 | 45.6 | — | |
| ShareGPT4Video2025.01 | 43.6 | — | |
| ShareGPT4VideoModel Size=8B2025.11 | 43.6 | 37.9 | |
| VideoLLaMA2Training Data=same as Video-LLaVA [36]2025.01 | 42.7 | — | |
| VideoLLaMBSize=7B2026.03 | 41.4 | — | |
| Video-LLaVATraining Data=same as Video-LLaVA [36]2025.01 | 40.4 | — | |
| Video-LLaVAModel Size=7B2025.11 | 40.4 | 38.1 | |
| Video-LLaVASize=7B, Evaluation Protocol=Standard2025.12 | 39.9 | — | |
| ShareGPT4Video-8BPerception budget constraint=false, Perc. budget=-2026.05 | 39.9 | — | |
| MovieChatSize=7B2026.03 | 38.2 | — | |
| LLaMA-VIDSize=7B2026.03 | 33.2 | — | |
| AdaSparkModel Size=Small Size, Backbone=Qwen-2.5-VL-3B, SFT=true2026.04 | — | 63.5 | |
| AdaSparkModel Size=Mid Size, Backbone=Qwen-2.5-VL-7B, SFT=true2026.04 | — | 66.2 | |
| FastVModel Size=Small Size, Backbone=Qwen-2.5-VL-3B, SFT=true2026.04 | — | 62.4 |