Video Understanding on Aggregate MVBench, LongVideo Bench, MLVU, VideoMME
100Average AccuracyQwen3-VL-8B-Instruct
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen3-VL-8B-InstructMax Input Frames=64, Retention Ratio=100.0%2026.03 | 100 | 67 | |
| Qwen3-VL-8B-InstructMax Input Frames=32, Retention Ratio=100.0%2026.03 | 100 | 64.2 | |
| Qwen3-VL-8B-InstructMax Input Frames=642026.03 | 100 | 67 | |
| Qwen3-VL-8B-InstructMax Input Frames=322026.03 | 100 | 64.2 | |
| V-CASTMax Input Frames=32, Retention Ratio=35%2026.03 | 98.8 | 63.5 | |
| V-CASTMax Input Frames=32, Retention Ratio=25%2026.03 | 98.6 | 63.3 | |
| FlashVIDMax Input Frames=32, Retention Ratio=35%2026.03 | 98.3 | 63.1 | |
| V-CASTMax Input Frames=64, Retention Ratio=35%2026.03 | 98.2 | 65.8 | |
| VidCom2Max Input Frames=32, Retention Ratio=35%2026.03 | 98 | 62.9 | |
| VidCom2Max Input Frames=64, Retention Ratio=35%2026.03 | 97.6 | 65.4 | |
| FlashVIDMax Input Frames=32, Retention Ratio=25%2026.03 | 97.5 | 62.6 | |
| FastVIDMax Input Frames=32, Retention Ratio=35%2026.03 | 97.5 | 62.6 | |
| V-CASTMax Input Frames=64, Retention Ratio=25%2026.03 | 97 | 65 | |
| VidCom2Max Input Frames=32, Retention Ratio=25%2026.03 | 96.6 | 62 | |
| FastVIDMax Input Frames=32, Retention Ratio=25%2026.03 | 96.3 | 61.8 | |
| V-CASTMax Input Frames=32, Retention Ratio=15%2026.03 | 95.8 | 61.5 | |
| HoliTomMax Input Frames=32, Retention Ratio=35%2026.03 | 95.8 | 61.5 | |
| FastVIDMax Input Frames=64, Retention Ratio=35%2026.03 | 95.7 | 64.1 | |
| FlashVIDMax Input Frames=32, Retention Ratio=15%2026.03 | 95.6 | 61.4 | |
| VidCom2Max Input Frames=64, Retention Ratio=25%2026.03 | 95.5 | 64 | |
| HoliTomMax Input Frames=64, Retention Ratio=35%2026.03 | 95 | 63.6 | |
| FastVIDMax Input Frames=64, Retention Ratio=25%2026.03 | 94.9 | 63.6 | |
| V-CASTMax Input Frames=64, Retention Ratio=15%2026.03 | 94.9 | 63.6 | |
| VisionZipMax Input Frames=32, Retention Ratio=35%2026.03 | 94.8 | 60.8 | |
| VisionZipMax Input Frames=64, Retention Ratio=35%2026.03 | 94.5 | 63.3 | |
| HoliTomMax Input Frames=32, Retention Ratio=25%2026.03 | 93.8 | 60.2 | |
| FastVIDMax Input Frames=32, Retention Ratio=15%2026.03 | 93.6 | 60.1 | |
| VisionZipMax Input Frames=32, Retention Ratio=25%2026.03 | 93.5 | 60 | |
| VisionZipMax Input Frames=64, Retention Ratio=25%2026.03 | 93.4 | 62.6 | |
| HoliTomMax Input Frames=64, Retention Ratio=25%2026.03 | 93.3 | 62.5 | |
| FastVIDMax Input Frames=64, Retention Ratio=15%2026.03 | 92.7 | 62.1 | |
| VidCom2Max Input Frames=32, Retention Ratio=15%2026.03 | 92.5 | 59.4 | |
| VidCom2Max Input Frames=64, Retention Ratio=15%2026.03 | 91.3 | 61.2 | |
| VisionZipMax Input Frames=32, Retention Ratio=15%2026.03 | 91.3 | 58.6 | |
| HoliTomMax Input Frames=32, Retention Ratio=15%2026.03 | 91 | 58.4 | |
| HoliTomMax Input Frames=64, Retention Ratio=15%2026.03 | 90.9 | 60.9 | |
| VisionZipMax Input Frames=64, Retention Ratio=15%2026.03 | 90.1 | 60.4 | |
| LLaVA-Video 7BRetention Ratio R=100%, # Newline Tokens M=8322026.05 | 63.6 | — | |
| LLaVA-Video-7BBase Model=LLaVA-Video-7B, Input Frames=642026.04 | 63.3 | 100 | |
| Qwen2.5-VL-7BBase Model=Qwen2.5-VL-7B, Input Frames=322026.04 | 63.2 | 100 | |
| Tango†Base Model=LLaVA-Video-7B, Retention Ratio=20%, Intra-LLM=true, Input Frames=642026.04 | 61.7 | 97.6 | |
| TangoBase Model=LLaVA-Video-7B, Retention Ratio=20%, Input Frames=642026.04 | 61.2 | 96.8 | |
| DyCokeRetention Ratio R=32.1%, # Newline Tokens M=2562026.05 | 61.1 | — | |
| FastVIDBase Model=LLaVA-Video-7B, Retention Ratio=20%, Input Frames=642026.04 | 60.9 | 96.3 | |
| HoliTomBase Model=LLaVA-Video-7B, Retention Ratio=20%, Input Frames=642026.04 | 60.8 | 96.1 | |
| DynaTokRetention Ratio R=24.9%, # Newline Tokens M=8322026.05 | 60.6 | — | |
| FastVRetention Ratio R=100%/25%, # Newline Tokens M=832/568.72026.05 | 60.3 | — | |
| Tango†Base Model=LLaVA-Video-7B, Retention Ratio=10%, Intra-LLM=true, Input Frames=642026.04 | 60.2 | 95.1 | |
| VisionZipBase Model=LLaVA-Video-7B, Retention Ratio=20%, Input Frames=642026.04 | 60.2 | 95.2 | |
| TangoBase Model=Qwen2.5-VL-7B, Retention Ratio=20%, Input Frames=322026.04 | 59.8 | 94.6 | |
| TangoBase Model=LLaVA-Video-7B, Retention Ratio=10%, Input Frames=642026.04 | 59.6 | 94.2 | |
| HoliTomBase Model=LLaVA-Video-7B, Retention Ratio=10%, Input Frames=642026.04 | 59.5 | 94.1 | |
| VisionZipBase Model=Qwen2.5-VL-7B, Retention Ratio=20%, Input Frames=322026.04 | 59.2 | 93.6 | |
| FastVIDBase Model=LLaVA-Video-7B, Retention Ratio=10%, Input Frames=642026.04 | 59 | 93.3 | |
| FastVIDBase Model=Qwen2.5-VL-7B, Retention Ratio=20%, Input Frames=322026.04 | 59 | 93.3 | |
| DynaTokRetention Ratio R=9.5%, # Newline Tokens M=8322026.05 | 58.3 | — | |
| VisionZipRetention Ratio R=24.9%, # Newline Tokens M=642026.05 | 57.8 | — | |
| FastVRetention Ratio R=100%/10%, # Newline Tokens M=832/311.82026.05 | 57 | — | |
| VisionZipBase Model=LLaVA-Video-7B, Retention Ratio=10%, Input Frames=642026.04 | 56.4 | 89.1 | |
| VidComBase Model=LLaVA-Video-7B, Retention Ratio=10%, Input Frames=642026.04 | 54.5 | 86.1 | |
| DyCokeRetention Ratio R=25%, # Newline Tokens M=2082026.05 | 54.2 | — | |
| FastVBase Model=LLaVA-Video-7B, Retention Ratio=10%, Input Frames=642026.04 | 48.9 | 77.3 | |
| VisionZipRetention Ratio R=9.5%, # Newline Tokens M=642026.05 | 48.7 | — | |
| FastVIDBase Model=Qwen3-VL-30B-A3B-Instruct, Retention Ratio=25%, Input Frames=642026.03 | — | 64 | |
| HoliTomBase Model=Qwen3-VL-30B-A3B-Instruct, Retention Ratio=25%, Input Frames=642026.03 | — | 64.8 | |
| V-CASTBase Model=Qwen3-VL-30B-A3B-Instruct, Retention Ratio=25%, Input Frames=642026.03 | — | 68.1 | |
| VidCom2Base Model=Qwen3-VL-30B-A3B-Instruct, Retention Ratio=25%, Input Frames=642026.03 | — | 68 | |
| VisionZipBase Model=Qwen3-VL-30B-A3B-Instruct, Retention Ratio=25%, Input Frames=642026.03 | — | 62.6 |