Video Understanding on Video-MME without subtitles
84.8Overall ScoreGemini-2.5-Pro
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| Gemini-2.5-Pro2026.05 | 84.8 | — | — | — | — | — | — | |
| STEMO-TrackBackbone=Gemini-3-Flash2026.05 | 76.8 | 85.8 | 78.3 | 66.2 | — | — | — | |
| Gemini-1.5-Pro2026.02 | 75 | 81.7 | 74.3 | 67.4 | — | — | — | |
| Gemini 1.5 ProSize (LLM)=-, #Visual Tokens=-2026.03 | 75 | — | — | 67.4 | — | — | — | |
| Gemini-1.5-Pro2026.05 | 75 | 81.7 | 74.3 | 67.4 | — | — | — | |
| Gemini-1.5-Pro# Frame=1 fps, Paradigm=Standard (w/o Tools), Adaptive keyframe retrieval=false2026.06 | 75 | 81.7 | 74.3 | 67.4 | — | — | — | |
| STEMO-TrackBackbone=Qwen3-VL-235B2026.05 | 74.1 | 82.2 | 76 | 62.7 | — | — | — | |
| Qwen2.5-VL-72BParams=72B2026.05 | 73.3 | — | — | — | — | — | — | |
| GPT-4o2026.02 | 71.9 | 80 | 70.3 | 65.3 | — | — | — | |
| GPT-4oSize (LLM)=-, #Visual Tokens=-2026.03 | 71.9 | — | — | 65.3 | — | — | — | |
| GPT-4o2026.05 | 71.9 | — | — | — | — | — | — | |
| GPT-4o2026.05 | 71.9 | 80 | 70.3 | 65.3 | — | — | — | |
| GPT-4o# Frame=384, Paradigm=Standard (w/o Tools), Adaptive keyframe retrieval=false2026.06 | 71.9 | 80 | 70.3 | 65.3 | — | — | — | |
| VINOType=Und. and Gen., Params=4B2026.01 | 69.3 | — | — | — | — | — | — | |
| InternVL3-8B (pub.)Params=8B, Ctx (k)=44.72026.05 | 66.3 | — | — | — | — | — | — | |
| LLaVA-OV2026.05 | 66.3 | 76.7 | 62.2 | 60 | — | — | — | |
| VideoLLaMA3Type=Und. only, Params=7B2026.01 | 66.2 | — | — | — | — | — | — | |
| InternVL3.5Type=Und. only, Params=4B2026.01 | 65.4 | — | — | — | — | — | — | |
| FlashAttentionBackbone=Qwen2.5-VL-7B-Instruct, Sparsity(↑)=0%2025.12 | 65.3 | 75.1 | 66.7 | 54.2 | — | — | — | |
| Qwen2.5-VLType=Und. only, Params=7B2026.01 | 65.1 | — | — | — | — | — | — | |
| MetaQuery-XLType=Und. and Gen., Params=7B2026.01 | 65.1 | — | — | — | — | — | — | |
| FlexPrefillBackbone=Qwen2.5-VL-7B-Instruct, Sparsity(↑)=27.3%, tau=0.1, gamma=0.992025.12 | 65 | 74.6 | 66.8 | 53.7 | — | — | — | |
| UniSparseBackbone=Qwen2.5-VL-7B-Instruct, Sparsity(↑)=45.9%, target_sparsity=0.952025.12 | 65 | 74.8 | 65.8 | 54.3 | — | — | — | |
| HIMMEL (Qwen3-VL)Params=8B, Ctx (k)=16.2, × speedup=2.8×2026.05 | 64.9 | — | — | — | — | — | — | |
| UniSparseBackbone=Qwen2.5-VL-7B-Instruct, Sparsity(↑)=57.7%, target_sparsity=0.92025.12 | 64.6 | 74.4 | 65.9 | 53.3 | — | — | — | |
| VideoTemp-o3-7B-RLParameters=7B2026.02 | 64.5 | 72.2 | 66.6 | 54.7 | — | — | — | |
| FlexPrefillBackbone=Qwen2.5-VL-7B-Instruct, Sparsity(↑)=57.9%, tau=0.1, gamma=0.952025.12 | 64.4 | 73.9 | 65.8 | 53.7 | — | — | — | |
| XAttentionBackbone=Qwen2.5-VL-7B-Instruct, Sparsity(↑)=49.2%, tau=0.952025.12 | 63.9 | 74.1 | 64.7 | 52.8 | — | — | — | |
| HIMMEL (Qwen2.5-VL)Params=7B, Ctx (k)=16.2, × speedup=2.8×2026.05 | 63.5 | — | — | — | — | — | — | |
| Reflect-R1# Frame=768, Paradigm=Tool-Augmented Reasoning, Adaptive keyframe retrieval=false2026.06 | 63.5 | 73.9 | 61 | 55.6 | — | — | — | |
| QViC-MFSize (LLM)=7B (Qwen2), #Visual Tokens=16, Sampling Rate=2 fps2026.03 | 63.4 | — | — | 54 | — | — | — | |
| LLaVA-VideoLLM=Qwen2-7B, Vision Encoder=SigLIP-SO400M2024.12 | 63.3 | — | — | — | — | — | — | |
| Qwen2-VL-InstructLLM=Qwen2-7B, Vision Encoder=DFN-CLIP-H2024.12 | 63.3 | — | — | — | — | — | — | |
| LLaVA-VideoSize (LLM)=7B (Qwen2), #Visual Tokens=1692026.03 | 63.3 | — | — | — | — | — | — | |
| LLaVA-VideoParams=7B, Ctx (k)=32.02026.05 | 63.3 | — | — | — | — | — | — | |
| XAttentionBackbone=Qwen2.5-VL-7B-Instruct, Sparsity(↑)=60.2%, tau=0.92025.12 | 63 | 73.6 | 63.6 | 51.9 | — | — | — | |
| TimeSearch-R-7B# Frame=768, Paradigm=Tool-Augmented Reasoning, Adaptive keyframe retrieval=false2026.06 | 62.7 | 73.5 | 61.2 | 53.4 | — | — | — | |
| QViC-MFSize (LLM)=7B (Qwen2), #Visual Tokens=16, Sampling Rate=1 fps2026.03 | 62.4 | — | — | 52.6 | — | — | — | |
| CAREFrames=642026.06 | 62.3 | — | — | — | — | — | — | |
| VideoChat-R1-7BParameters=7B, reproduced=true2026.02 | 62.1 | 72.2 | 60.7 | 50.9 | — | — | — | |
| TASMContext Length=2.3k, Peak Mem=24G2026.06 | 61.5 | 68.6 | 60.9 | 55.1 | — | — | — | |
| T* (GPT-4o)# Frame=32†, Paradigm=Tool-Augmented Reasoning, Adaptive keyframe retrieval=true2026.06 | 61.45 | 72.1 | 60.3 | 52 | — | — | — | |
| Video-R1-7BParameters=7B, reproduced=true2026.02 | 61.4 | 74.1 | 61.1 | 51.2 | — | — | — | |
| Vision-R1-7B# Frame=768, Paradigm=Standard (w/o Tools), Adaptive keyframe retrieval=false2026.06 | 61.4 | 71.2 | 60.2 | 52.8 | — | — | — | |
| Qwen2.5-VL-7B-Instruct# Frame=768, Paradigm=Standard (w/o Tools), Adaptive keyframe retrieval=false2026.06 | 61.3 | 71.4 | 60.1 | 52.3 | — | — | — | |
| VL-Rethinker# Frame=768, Paradigm=Standard (w/o Tools), Adaptive keyframe retrieval=false2026.06 | 61.3 | 70.1 | 60.8 | 53 | — | — | — | |
| Flash-VStreamSize (LLM)=7B (Qwen2), #Visual Tokens=1282026.03 | 61.2 | — | — | 50.3 | — | — | — | |
| Qwen2.5-VL (32 fr.)Params=7B, Ctx (k)=44.7, × speedup=1.0×2026.05 | 61.2 | — | — | — | — | — | — | |
| Video-LLaVAParams=7B, Ctx (k)=8.02026.05 | 60.9 | — | — | — | — | — | — | |
| Video-R1-7B# Frame=768, Paradigm=Standard (w/o Tools), Adaptive keyframe retrieval=false2026.06 | 60.8 | 72.2 | 58.1 | 52.3 | — | — | — | |
| VideoTemp-o3-7B-SFTParameters=7B2026.02 | 60.6 | 72 | 59.2 | 50.2 | — | — | — | |
| LongVUSize (LLM)=7B (Qwen2), #Visual Tokens=642026.03 | 60.6 | — | — | — | — | — | — | |
| CAREFrames=322026.06 | 60.6 | — | — | — | — | — | — | |
| MLoCContext Length=27.9k, Peak Mem=38G2026.06 | 60.3 | 68.5 | 60.2 | 53.4 | — | — | — | |
| VILA-1.5-40BParameters=40B2026.02 | 60.1 | — | — | — | — | — | — | |
| EMLoCContext Length=2.3k, Peak Mem=24G2026.06 | 60.1 | 68.2 | 59.8 | 49.5 | — | — | — | |
| VILA-1.5-40B2026.06 | 60.1 | — | — | — | — | — | — | |
| LLaVA-Video-7BVisual Token Budget=All 2,704 Visual Tokens, Evaluation Mode=16-frame evaluation, Base Model=LLaVA-Video-7B2026.07 | 59.96 | — | — | — | — | — | 100 | |
| Qwen2.5-VL-7BParameters=7B, reproduced=true2026.02 | 59.9 | 69.8 | 59.2 | 50.8 | — | — | — | |
| StreamingTOMFr. (Sampling Rate/Frames)=0.5/0.2fps, Design Category=Streaming-Designed, Training Protocol=Training-free2025.10 | 59.9 | 71.3 | 57.8 | 50.6 | — | — | — | |
| VideoRFT-7BParameters=7B2026.02 | 59.8 | — | — | — | — | — | — | |
| LLaVA-OV-7B + StreamMemFr. (Sampling Rate/Frames)=0.5/0.2fps, Design Category=Streaming-Designed, Training Protocol=Training-free2025.10 | 59.4 | 71.5 | 56.6 | 50.1 | — | — | — | |
| FastVVisual Token Budget=Retain 1,024 Tokens, Evaluation Mode=16-frame evaluation, Base Model=LLaVA-Video-7B, Source=ECCV’242026.07 | 59.26 | — | — | — | — | — | 97.7 | |
| VILA-1.5LLM Params=34B, Frames=82024.06 | 59 | 68.1 | 58.1 | 50.8 | — | — | — | |
| Video-MTR-7BParameters=7B2026.02 | 59 | — | — | 51 | — | — | — | |
| LLaVA-OV-7B + HoliTomFr. (Sampling Rate/Frames)=32, Design Category=Offline-Designed, Training Protocol=Training-free2025.10 | 58.9 | — | — | — | — | — | — | |
| LLaVA-OneVisionLLM Size=7B, Frames=322025.12 | 58.5 | 70.1 | 56.4 | 48.9 | — | — | — | |
| LLaVA-OV-7BFr. (Sampling Rate/Frames)=32, Design Category=Offline-Designed, Training Protocol=Training-free2025.10 | 58.4 | 70.1 | 56.4 | 48.8 | — | — | — | |
| LLaVA-OV-7BFr. (Sampling Rate/Frames)=32, Design Category=Streaming-Designed, Training Protocol=Training-free2025.10 | 58.4 | 70.1 | 56.4 | 48.8 | — | — | — | |
| LLAVA-OVLLM=Qwen2-7B, Vision Encoder=SigLIP-SO400M2024.12 | 58.2 | — | — | — | — | — | — | |
| LLaVA-OV-7B + VisionZipFr. (Sampling Rate/Frames)=32, Design Category=Offline-Designed, Training Protocol=Training-free2025.10 | 58.2 | — | — | — | — | — | — | |
| LLaVA-OneVision-7B2026.06 | 58.2 | — | — | — | — | — | — | |
| AnchorPruneVisual Token Budget=Retain 1,024 Tokens, Evaluation Mode=16-frame evaluation, Base Model=LLaVA-Video-7B2026.07 | 58.19 | — | — | — | — | — | 98.3 | |
| CDPrunerVisual Token Budget=Retain 1,024 Tokens, Evaluation Mode=16-frame evaluation, Base Model=LLaVA-Video-7B, Source=NeurIPS’252026.07 | 57.85 | — | — | — | — | — | 96.8 | |
| Temporal-RLTFrames=322026.06 | 57.6 | — | — | — | — | — | — | |
| Frame-VoyagerSize (LLM)=7B (Qwen2), #Visual Tokens=1962026.03 | 57.5 | — | — | 48.9 | — | — | — | |
| show-o2Type=Und. and Gen., Params=7B2026.01 | 57.4 | — | — | — | — | — | — | |
| TimeMaker-8BParameters=8B2026.02 | 57.3 | — | — | 46.4 | — | — | — | |
| LLaVA-OV-7B + LiveVLMFr. (Sampling Rate/Frames)=0.5/0.2fps, Design Category=Streaming-Designed, Training Protocol=Training-free2025.10 | 57.3 | 66.7 | 56.4 | 48.8 | — | — | — | |
| CAREFrames=162026.06 | 57.3 | — | — | — | — | — | — | |
| DivPruneVisual Token Budget=Retain 1,024 Tokens, Evaluation Mode=16-frame evaluation, Base Model=LLaVA-Video-7B, Source=CVPR’252026.07 | 57 | — | — | — | — | — | 96.1 | |
| Dispider-7BFr. (Sampling Rate/Frames)=1fps, Design Category=Streaming-Designed, Training Protocol=Training-based2025.10 | 56.5 | — | 53.7 | 49.7 | — | — | — | |
| FastVVisual Token Budget=Retain 512 Tokens, Evaluation Mode=16-frame evaluation, Base Model=LLaVA-Video-7B, Source=ECCV’242026.07 | 56.37 | — | — | — | — | — | 92.4 | |
| AnchorPruneVisual Token Budget=Retain 512 Tokens, Evaluation Mode=16-frame evaluation, Base Model=LLaVA-Video-7B2026.07 | 56.26 | — | — | — | — | — | 94.1 | |
| Kangeroo-8B2026.06 | 56 | — | — | — | — | — | — | |
| VideoAgent (GPT-4)# Frame=87†, Paradigm=Tool-Augmented Reasoning, Adaptive keyframe retrieval=true2026.06 | 56 | — | — | 49 | — | — | — | |
| DivPruneVisual Token Budget=Retain 512 Tokens, Evaluation Mode=16-frame evaluation, Base Model=LLaVA-Video-7B, Source=CVPR’252026.07 | 55.89 | — | — | — | — | — | 93.4 | |
| VideoXL-7BParameters=7B2026.02 | 55.5 | — | — | — | — | — | — | |
| Video-XLSize (LLM)=7B (Qwen2), #Visual Tokens=162026.03 | 55.5 | — | — | — | — | — | — | |
| LLaVA-OV-7B + DyCokeFr. (Sampling Rate/Frames)=32, Design Category=Offline-Designed, Training Protocol=Training-free2025.10 | 54.3 | — | — | — | — | — | — | |
| Video-R1-7BFrames=162026.06 | 54.3 | — | — | — | — | — | — | |
| CDPrunerVisual Token Budget=Retain 512 Tokens, Evaluation Mode=16-frame evaluation, Base Model=LLaVA-Video-7B, Source=NeurIPS’252026.07 | 54.22 | — | — | — | — | — | 91.7 | |
| VideoChat2Params=7B, Ctx (k)=32.02026.05 | 54.1 | — | — | — | — | — | — | |
| InternVL2LLM=InternLM-7B, Vision Encoder=InternViT-300M2024.12 | 54 | — | — | — | — | — | — | |
| LLaVA-NeXT-INST-ITLLM=Qwen2-7B, Vision Encoder=SigLIP-SO4002024.12 | 54 | — | — | — | — | — | — | |
| Qwen2-VL + SF2TBackbone=Qwen2-VL, Training Configuration=Base+SF2T2025.04 | 53.6 | — | — | — | — | — | — | |
| MiniCPM-V 2.6 + SF2TBackbone=MiniCPM-V 2.6, Training Configuration=Base+SF2T2025.04 | 53.19 | — | — | — | — | — | — | |
| Qwen2.5-VL-7B2026.06 | 52.8 | — | — | — | — | — | — | |
| LongVALLM Params=7B, Frames=128, Zero-Shot=true2024.06 | 52.6 | 61.1 | 50.4 | 46.2 | — | — | — | |
| LongVA-7BParameters=7B2026.02 | 52.6 | 61.1 | 50.4 | 46.2 | — | — | — |