Video Understanding on MLVU (test)
100.3AverageFSR
Evaluation Results
| Method | Links | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| FSRReduction Ratio=50%, Frames=322026.02 | 100.3 | 0.502 | — | — | — | — | — | — | — | — | — | |
| LLaVA-Video-7B-qwen2Reduction Ratio=100%, Frames=322026.02 | 100 | 0.501 | — | — | — | — | — | — | — | — | — | |
| FSRReduction Ratio=60%, Frames=322026.02 | 99.6 | 0.5 | — | — | — | — | — | — | — | — | — | |
| HoloVReduction Ratio=50%, Frames=322026.02 | 99.2 | 0.491 | — | — | — | — | — | — | — | — | — | |
| FSRReduction Ratio=70%, Frames=322026.02 | 98.9 | 0.476 | — | — | — | — | — | — | — | — | — | |
| HoloVReduction Ratio=60%, Frames=322026.02 | 98.5 | 0.491 | — | — | — | — | — | — | — | — | — | |
| HoloVReduction Ratio=70%, Frames=322026.02 | 98.2 | 0.485 | — | — | — | — | — | — | — | — | — | |
| FSRReduction Ratio=80%, Frames=322026.02 | 98.2 | 0.465 | — | — | — | — | — | — | — | — | — | |
| HoloVReduction Ratio=80%, Frames=322026.02 | 98 | 0.465 | — | — | — | — | — | — | — | — | — | |
| InternVL3 + HFSLLM size=8B+1.5B, # Frames=162025.12 | 50 | — | — | — | — | — | — | — | — | — | — | |
| InternVL3LLM size=8B, # Frames=162025.12 | 46 | — | — | — | — | — | — | — | — | — | — | |
| InternVL2LLM Size=76B, Vision Encoder=InternViT-6B, Training Regime=Training-Based2026.02 | 45.7 | — | 85.7 | 51.3 | 48.3 | 47.2 | 52 | 44.4 | 32.9 | 15 | 34.9 | |
| Video-LLaMA2LLM Size=72B, Vision Encoder=CLIP-L, Training Regime=Training-Based2026.02 | 45.6 | — | 80.2 | 53.8 | 36.7 | 54.7 | 54 | 38.9 | 42.9 | 16.7 | 32.6 | |
| VideoLLaMA2LLM size=7B, # Frames=82025.12 | 45.6 | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL + HFSLLM size=7B+1.5B, # Frames=162025.12 | 45.6 | — | — | — | — | — | — | — | — | — | — | |
| Video-XLLLM size=7B, # Frames=1282025.12 | 45.5 | — | — | — | — | — | — | — | — | — | — | |
| KTV-34B-denseLLM Size=34B, Vision Encoder=CLIP-L, Training Regime=Training-Free2026.02 | 45 | — | 85.7 | 51.3 | 48.7 | 47.2 | 48 | 58.3 | 35.7 | 10 | 37.2 | |
| KTV-34B-sparseLLM Size=34B, Vision Encoder=CLIP-L, Training Regime=Training-Free2026.02 | 44.8 | — | 81.3 | 51.3 | 53.3 | 47.2 | 50 | 52.8 | 37.1 | 11.7 | 34.9 | |
| KTV-34B-normalLLM Size=34B, Vision Encoder=CLIP-L, Training Regime=Training-Free2026.02 | 44.2 | — | 81.3 | 56.4 | 46.7 | 45.3 | 46 | 52.8 | 32.9 | 8.3 | 41.9 | |
| VILA-1.5LLM size=40B, # Frames=142025.12 | 44.2 | — | — | — | — | — | — | — | — | — | — | |
| SF-LLaVA-34BLLM Size=34B, Vision Encoder=CLIP-L, Training Regime=Training-Free2026.02 | 43.6 | — | 76.9 | 43.6 | 36.7 | 39.6 | 44 | 47.2 | 27.1 | 8.3 | 32.6 | |
| Video-CCAMLLM size=14B, # Frames=962025.12 | 42.9 | — | — | — | — | — | — | — | — | — | — | |
| Qwen2.5-VLLLM size=7B, # Frames=162025.12 | 41.8 | — | — | — | — | — | — | — | — | — | — | |
| LongVALLM size=7B, # Frames=128/2562025.12 | 41.1 | — | — | — | — | — | — | — | — | — | — | |
| KTV-7B-denseLLM Size=7B, Vision Encoder=CLIP-L, Training Regime=Training-Free2026.02 | 36.9 | — | 69.2 | 48.9 | 33.3 | 39.6 | 34 | 38.9 | 27.1 | 23.3 | 27.9 | |
| KTV-7B-sparseLLM Size=7B, Vision Encoder=CLIP-L, Training Regime=Training-Free2026.02 | 36.5 | — | 73.6 | 43.6 | 35 | 41.5 | 34 | 36.1 | 27.1 | 18.3 | 20.9 | |
| KTV-7B-normalLLM Size=7B, Vision Encoder=CLIP-L, Training Regime=Training-Free2026.02 | 36.1 | — | 72.5 | 51.3 | 35 | 41.5 | 34 | 38.9 | 24.3 | 21.7 | 25.6 | |
| ShareGPT4VideoLLM size=8B, # Frames=162025.12 | 33.8 | — | — | — | — | — | — | — | — | — | — | |
| IG-VLM-34BLLM Size=34B, Vision Encoder=CLIP-L, Training Regime=Training-Free2026.02 | 33.3 | — | 68.1 | 35.9 | 21.7 | 30.1 | 34 | 50 | 20 | 6.7 | 20.1 | |
| SF-LLaVA-7BLLM Size=7B, Vision Encoder=CLIP-L, Training Regime=Training-Free2026.02 | 32.7 | — | 68.1 | 23.1 | 25 | 34 | 24 | 33.3 | 22.9 | 16.7 | 16.3 | |
| IG-VLM-7BLLM Size=7B, Vision Encoder=CLIP-L, Training Regime=Training-Free2026.02 | 32.1 | — | 74.7 | 33.3 | 26.7 | 18.9 | 24 | 33.3 | 17.1 | 15 | 27.9 | |
| Video-LLaVALLM Size=7B, Vision Encoder=CLIP-L, Training Regime=Training-Based2026.02 | 30.7 | — | 70.3 | 38.5 | 13.3 | 26.4 | 26 | 38.9 | 20 | 21.7 | 20.9 | |
| Video-LLaVALLM size=7B, # Frames=82025.12 | 30.7 | — | — | — | — | — | — | — | — | — | — | |
| Video-LLaMA2LLM Size=13B, Vision Encoder=CLIP-L, Training Regime=Training-Based2026.02 | 18.9 | — | 52.7 | 12.8 | 13.3 | 17 | 12 | 19.4 | 15.7 | 8.3 | 18.6 | |
| FastV (ECCV2024)Token Retention Ratio=75%, Backbone=Qwen2.5-VL, Input Frames=322025.12 | — | 0.415 | — | — | — | — | — | — | — | — | — | |
| FastV (ECCV2024)Token Retention Ratio=50%, Backbone=Qwen2.5-VL, Input Frames=322025.12 | — | 0.368 | — | — | — | — | — | — | — | — | — | |
| FastV (ECCV2024)Token Retention Ratio=25%, Backbone=Qwen2.5-VL, Input Frames=322025.12 | — | 0.34 | — | — | — | — | — | — | — | — | — | |
| FastV (ECCV2024)Token Retention Ratio=15%, Backbone=Qwen2.5-VL, Input Frames=322025.12 | — | 0.328 | — | — | — | — | — | — | — | — | — | |
| FastV (ECCV2024)Token Retention Ratio=10%, Backbone=Qwen2.5-VL, Input Frames=322025.12 | — | 0.246 | — | — | — | — | — | — | — | — | — | |
| FastV (ECCV2024) + DyToK (7B)Token Retention Ratio=75%, Backbone=Qwen2.5-VL, Input Frames=32, Model Scale=7B2025.12 | — | 0.454 | — | — | — | — | — | — | — | — | — | |
| FastV (ECCV2024) + DyToK (7B)Token Retention Ratio=50%, Backbone=Qwen2.5-VL, Input Frames=32, Model Scale=7B2025.12 | — | 0.441 | — | — | — | — | — | — | — | — | — | |
| FastV (ECCV2024) + DyToK (7B)Token Retention Ratio=25%, Backbone=Qwen2.5-VL, Input Frames=32, Model Scale=7B2025.12 | — | 0.433 | — | — | — | — | — | — | — | — | — | |
| FastV (ECCV2024) + DyToK (7B)Token Retention Ratio=15%, Backbone=Qwen2.5-VL, Input Frames=32, Model Scale=7B2025.12 | — | 0.398 | — | — | — | — | — | — | — | — | — | |
| FastV (ECCV2024) + DyToK (7B)Token Retention Ratio=10%, Backbone=Qwen2.5-VL, Input Frames=32, Model Scale=7B2025.12 | — | 0.246 | — | — | — | — | — | — | — | — | — | |
| Vanilla (TMLR)Token Retention Ratio=100%, Backbone=Qwen2.5-VL, Input Frames=322025.12 | — | 0.447 | — | — | — | — | — | — | — | — | — | |
| VisionZip†Token Retention Ratio=75%, Backbone=Qwen2.5-VL, Input Frames=32, Pooling Compatibility=true2025.12 | — | 0.446 | — | — | — | — | — | — | — | — | — | |
| VisionZip†Token Retention Ratio=50%, Backbone=Qwen2.5-VL, Input Frames=32, Pooling Compatibility=true2025.12 | — | 0.45 | — | — | — | — | — | — | — | — | — | |
| VisionZip†Token Retention Ratio=25%, Backbone=Qwen2.5-VL, Input Frames=32, Pooling Compatibility=true2025.12 | — | 0.443 | — | — | — | — | — | — | — | — | — | |
| VisionZip†Token Retention Ratio=15%, Backbone=Qwen2.5-VL, Input Frames=32, Pooling Compatibility=true2025.12 | — | 0.416 | — | — | — | — | — | — | — | — | — | |
| VisionZip†Token Retention Ratio=10%, Backbone=Qwen2.5-VL, Input Frames=32, Pooling Compatibility=true2025.12 | — | 0.417 | — | — | — | — | — | — | — | — | — | |
| VisionZip† + DyToK (7B)Token Retention Ratio=75%, Backbone=Qwen2.5-VL, Input Frames=32, Model Scale=7B, Pooling Compatibility=true2025.12 | — | 0.431 | — | — | — | — | — | — | — | — | — | |
| VisionZip† + DyToK (7B)Token Retention Ratio=50%, Backbone=Qwen2.5-VL, Input Frames=32, Model Scale=7B, Pooling Compatibility=true2025.12 | — | 0.438 | — | — | — | — | — | — | — | — | — | |
| VisionZip† + DyToK (7B)Token Retention Ratio=25%, Backbone=Qwen2.5-VL, Input Frames=32, Model Scale=7B, Pooling Compatibility=true2025.12 | — | 0.445 | — | — | — | — | — | — | — | — | — | |
| VisionZip† + DyToK (7B)Token Retention Ratio=15%, Backbone=Qwen2.5-VL, Input Frames=32, Model Scale=7B, Pooling Compatibility=true2025.12 | — | 0.425 | — | — | — | — | — | — | — | — | — | |
| VisionZip† + DyToK (7B)Token Retention Ratio=10%, Backbone=Qwen2.5-VL, Input Frames=32, Model Scale=7B, Pooling Compatibility=true2025.12 | — | 0.436 | — | — | — | — | — | — | — | — | — |