Video Understanding on LongVideoBench (test)
82.3Accuracy (Overall)VideoChat-M1
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| VideoChat-M1Frames=69.9, Inference Time=19.8s2025.11 | 82.3 | — | — | — | — | — | |
| GPT-4oFrames=384, Inference Time=153.6s2025.11 | 66.7 | — | — | — | — | — | |
| Gemini-1.5-ProFrames=568, Inference Time=227.2s2025.11 | 64 | — | — | — | — | — | |
| SPAVABackbone=Qwen2.5VL-7B, Performance Degradation=No2026.01 | 59.76 | 77.78 | 74.42 | 57.77 | 50.71 | — | |
| APBBackbone=Qwen2.5VL-7B, Performance Degradation=No2026.01 | 59.16 | 75.13 | 73.84 | 58.01 | 50.18 | — | |
| FULLATTNBackbone=Qwen2.5VL-7B2026.01 | 58.38 | 73.81 | 73.84 | 56.44 | 49.82 | — | |
| STARATTNBackbone=Qwen2.5VL-7B, Performance Degradation=Yes2026.01 | 58.26 | 74.07 | 72.67 | 57.28 | 50.53 | — | |
| XATTNBackbone=Qwen2.5VL-7B, Performance Degradation=Yes2026.01 | 57.59 | 72.49 | 73.84 | 57.77 | 47.52 | — | |
| SLOWFASTBackbone=Qwen2.5VL-7B, Performance Degradation=Yes2026.01 | 56.47 | 72.49 | 68.6 | 53.16 | 49.82 | — | |
| Qwen2-VL-72BFrames=568, Inference Time=90.5s2025.11 | 55.6 | — | — | — | — | — | |
| SPAVABackbone=InternVL3-2B, Performance Degradation=No2026.01 | 55.42 | 62.43 | 67.44 | 54.85 | 49.82 | — | |
| FULLATTNBackbone=InternVL3-2B2026.01 | 55.35 | 61.9 | 67.44 | 54.61 | 50 | — | |
| APBBackbone=InternVL3-2B, Performance Degradation=Yes2026.01 | 55.2 | 61.9 | 67.44 | 55.83 | 48.76 | — | |
| XATTNBackbone=InternVL3-2B, Performance Degradation=Yes2026.01 | 54.97 | 62.43 | 71.51 | 53.4 | 48.58 | — | |
| APBBackbone=Qwen2.5VL-3B, Performance Degradation=No2026.01 | 54.97 | 68.25 | 72.09 | 52.18 | 47.34 | — | |
| STARATTNBackbone=Qwen2.5VL-3B, Performance Degradation=No2026.01 | 54.52 | 68.78 | 70.93 | 53.64 | 45.39 | — | |
| STARATTNBackbone=InternVL3-2B, Performance Degradation=Yes2026.01 | 54.3 | 63.49 | 66.86 | 52.91 | 48.4 | — | |
| SPAVABackbone=Qwen2.5VL-3B, Performance Degradation=No2026.01 | 54.3 | 68.25 | 70.93 | 53.88 | 44.86 | — | |
| FULLATTNBackbone=Qwen2.5VL-3B2026.01 | 53.47 | 69.31 | 70.93 | 52.18 | 43.79 | — | |
| XATTNBackbone=Qwen2.5VL-3B, Performance Degradation=Yes2026.01 | 53.03 | 66.14 | 69.19 | 52.18 | 44.33 | — | |
| SLOWFASTBackbone=Qwen2.5VL-3B, Performance Degradation=Yes2026.01 | 52.73 | 65.61 | 64.53 | 50.49 | 46.45 | — | |
| SLOWFASTBackbone=InternVL3-2B, Performance Degradation=Yes2026.01 | 51.46 | 60.85 | 66.86 | 50.24 | 44.5 | — | |
| SPARGEBackbone=InternVL3-2B, Performance Degradation=Yes2026.01 | 49.74 | 56.08 | 64.53 | 50.49 | 42.55 | — | |
| SPARGEBackbone=Qwen2.5VL-7B, Performance Degradation=Yes2026.01 | 49.37 | 65.08 | 66.86 | 47.09 | 40.43 | — | |
| SPARGEBackbone=Qwen2.5VL-3B, Performance Degradation=Yes2026.01 | 47.87 | 59.79 | 60.47 | 42.48 | 43.97 | — | |
| FastV (ECCV2024)Token Retention Ratio=75%, Backbone=Qwen2.5-VL, Input Frames=322025.12 | — | — | — | — | — | 56.1 | |
| FastV (ECCV2024)Token Retention Ratio=50%, Backbone=Qwen2.5-VL, Input Frames=322025.12 | — | — | — | — | — | 53.2 | |
| FastV (ECCV2024)Token Retention Ratio=25%, Backbone=Qwen2.5-VL, Input Frames=322025.12 | — | — | — | — | — | 49.3 | |
| FastV (ECCV2024)Token Retention Ratio=15%, Backbone=Qwen2.5-VL, Input Frames=322025.12 | — | — | — | — | — | 47.1 | |
| FastV (ECCV2024)Token Retention Ratio=10%, Backbone=Qwen2.5-VL, Input Frames=322025.12 | — | — | — | — | — | 42.3 | |
| FastV (ECCV2024) + DyToK (7B)Token Retention Ratio=75%, Backbone=Qwen2.5-VL, Input Frames=32, Model Scale=7B2025.12 | — | — | — | — | — | 58.3 | |
| FastV (ECCV2024) + DyToK (7B)Token Retention Ratio=50%, Backbone=Qwen2.5-VL, Input Frames=32, Model Scale=7B2025.12 | — | — | — | — | — | 57.9 | |
| FastV (ECCV2024) + DyToK (7B)Token Retention Ratio=25%, Backbone=Qwen2.5-VL, Input Frames=32, Model Scale=7B2025.12 | — | — | — | — | — | 54.5 | |
| FastV (ECCV2024) + DyToK (7B)Token Retention Ratio=15%, Backbone=Qwen2.5-VL, Input Frames=32, Model Scale=7B2025.12 | — | — | — | — | — | 51.1 | |
| FastV (ECCV2024) + DyToK (7B)Token Retention Ratio=10%, Backbone=Qwen2.5-VL, Input Frames=32, Model Scale=7B2025.12 | — | — | — | — | — | 42.3 | |
| Vanilla (TMLR)Token Retention Ratio=100%, Backbone=Qwen2.5-VL, Input Frames=322025.12 | — | — | — | — | — | 57.7 | |
| VisionZip†Token Retention Ratio=75%, Backbone=Qwen2.5-VL, Input Frames=32, Pooling Compatibility=true2025.12 | — | — | — | — | — | 56.5 | |
| VisionZip†Token Retention Ratio=50%, Backbone=Qwen2.5-VL, Input Frames=32, Pooling Compatibility=true2025.12 | — | — | — | — | — | 56.8 | |
| VisionZip†Token Retention Ratio=25%, Backbone=Qwen2.5-VL, Input Frames=32, Pooling Compatibility=true2025.12 | — | — | — | — | — | 55.1 | |
| VisionZip†Token Retention Ratio=15%, Backbone=Qwen2.5-VL, Input Frames=32, Pooling Compatibility=true2025.12 | — | — | — | — | — | 55.4 | |
| VisionZip†Token Retention Ratio=10%, Backbone=Qwen2.5-VL, Input Frames=32, Pooling Compatibility=true2025.12 | — | — | — | — | — | 53.2 | |
| VisionZip† + DyToK (7B)Token Retention Ratio=75%, Backbone=Qwen2.5-VL, Input Frames=32, Model Scale=7B, Pooling Compatibility=true2025.12 | — | — | — | — | — | 57 | |
| VisionZip† + DyToK (7B)Token Retention Ratio=50%, Backbone=Qwen2.5-VL, Input Frames=32, Model Scale=7B, Pooling Compatibility=true2025.12 | — | — | — | — | — | 56.8 | |
| VisionZip† + DyToK (7B)Token Retention Ratio=25%, Backbone=Qwen2.5-VL, Input Frames=32, Model Scale=7B, Pooling Compatibility=true2025.12 | — | — | — | — | — | 55.3 | |
| VisionZip† + DyToK (7B)Token Retention Ratio=15%, Backbone=Qwen2.5-VL, Input Frames=32, Model Scale=7B, Pooling Compatibility=true2025.12 | — | — | — | — | — | 54.1 | |
| VisionZip† + DyToK (7B)Token Retention Ratio=10%, Backbone=Qwen2.5-VL, Input Frames=32, Model Scale=7B, Pooling Compatibility=true2025.12 | — | — | — | — | — | 53.1 |