Video Multimodal Understanding on Video MME
61.9ScoreThinkStream-3B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ThinkStream-3BModel=ThinkStream-3B2026.03 | 61.9 | — | |
| Streamo-3BModel=Streamo-3B2026.03 | 61.8 | — | |
| Qwen2.5-VL-3BModel=Qwen2.5-VL-3B2026.03 | 61.5 | — | |
| Flash-VStream-7BModel=Flash-VStream-7B2026.03 | 61.2 | — | |
| Full ModelSparsity Ratio=0%, Base Model=LLaVA-OneVision2025.04 | 60.1 | — | |
| LLaVA-OV-7BPrefilling FLOPs (T)=40.8, FLOPs Ratio=100%, Before LLM Retained Ratio=100%2026.03 | 58.6 | — | |
| VisionZipPrefilling FLOPs (T)=8.7, FLOPs Ratio=21.3%, Before LLM Retained Ratio=25%2026.03 | 58.2 | — | |
| FastVIDPrefilling FLOPs (T)=8.7, FLOPs Ratio=21.3%, Before LLM Retained Ratio=25%2026.03 | 58 | — | |
| VisionZipPrefilling FLOPs (T)=7.0, FLOPs Ratio=17.2%, Before LLM Retained Ratio=20%2026.03 | 57.9 | — | |
| FastVIDPrefilling FLOPs (T)=7.0, FLOPs Ratio=17.2%, Before LLM Retained Ratio=20%2026.03 | 57.9 | — | |
| FastVIDPrefilling FLOPs (T)=5.2, FLOPs Ratio=12.7%, Before LLM Retained Ratio=15%2026.03 | 57.7 | — | |
| PruneVidPrefilling FLOPs (T)=8.7, FLOPs Ratio=21.3%, Before LLM Retained Ratio=25%2026.03 | 57.4 | — | |
| FastVIDPrefilling FLOPs (T)=3.4, FLOPs Ratio=8.3%, Before LLM Retained Ratio=10%2026.03 | 57.3 | — | |
| Dispider-7BModel=Dispider-7B2026.03 | 57.2 | — | |
| AOTPrefilling FLOPs (T)=8.7, FLOPs Ratio=21.3%, Before LLM Retained Ratio=25%2026.03 | 56.9 | — | |
| PruneVidPrefilling FLOPs (T)=7.0, FLOPs Ratio=17.2%, Before LLM Retained Ratio=20%2026.03 | 56.9 | — | |
| AOTPrefilling FLOPs (T)=7.0, FLOPs Ratio=17.2%, Before LLM Retained Ratio=20%2026.03 | 56.8 | — | |
| PruneVidPrefilling FLOPs (T)=5.2, FLOPs Ratio=12.7%, Before LLM Retained Ratio=15%2026.03 | 56.6 | — | |
| AOTPrefilling FLOPs (T)=3.4, FLOPs Ratio=8.3%, Before LLM Retained Ratio=10%2026.03 | 56.6 | — | |
| PDropPrefilling FLOPs (T)=10.5, FLOPs Ratio=25.7%, Before LLM Retained Ratio=100%2026.03 | 56.4 | — | |
| FastVPrefilling FLOPs (T)=9.3, FLOPs Ratio=22.8%, Before LLM Retained Ratio=100%2026.03 | 56.1 | — | |
| VisionZipPrefilling FLOPs (T)=5.2, FLOPs Ratio=12.7%, Before LLM Retained Ratio=15%2026.03 | 56.1 | — | |
| AOTPrefilling FLOPs (T)=5.2, FLOPs Ratio=12.7%, Before LLM Retained Ratio=15%2026.03 | 56.1 | — | |
| PruneVidPrefilling FLOPs (T)=3.4, FLOPs Ratio=8.3%, Before LLM Retained Ratio=10%2026.03 | 56 | — | |
| SparseGPTSparsity Ratio=60%, Base Model=LLaVA-OneVision2025.04 | 55.2 | — | |
| ECOFLaPSparsity Ratio=60%, Base Model=LLaVA-OneVision2025.04 | 54.4 | — | |
| DyCokePrefilling FLOPs (T)=8.7, FLOPs Ratio=21.3%, Before LLM Retained Ratio=25%2026.03 | 54.3 | — | |
| TAMPSparsity Ratio=60%, Base Model=LLaVA-OneVision2025.04 | 54 | — | |
| VisionZipPrefilling FLOPs (T)=3.4, FLOPs Ratio=8.3%, Before LLM Retained Ratio=10%2026.03 | 53.4 | — | |
| WandaSparsity Ratio=60%, Base Model=LLaVA-OneVision2025.04 | 51.5 | — | |
| OWLSparsity Ratio=60%, Base Model=LLaVA-OneVision2025.04 | 50.8 | — | |
| Full ModelBackbone=VideoLLaMA2 (7B), Sparsity Ratio=0%, Evaluation Protocol=Zero-shot2026.04 | 48.7 | 100 | |
| TopoVLM (Ours)Backbone=VideoLLaMA2 (7B), Sparsity Ratio=60%, Evaluation Protocol=Zero-shot2026.04 | 48 | 96.7 | |
| LLM-StreamlineBackbone=VideoLLaMA2 (7B), Sparsity Ratio=60%, Evaluation Protocol=Zero-shot2026.04 | 46.1 | 86.5 | |
| TAMPBackbone=VideoLLaMA2 (7B), Sparsity Ratio=60%, Evaluation Protocol=Zero-shot2026.04 | 42.5 | 95 | |
| LLM-PrunerBackbone=VideoLLaMA2 (7B), Sparsity Ratio=60%, Evaluation Protocol=Zero-shot2026.04 | 41.5 | 82.1 | |
| ECoFLaPBackbone=VideoLLaMA2 (7B), Sparsity Ratio=60%, Evaluation Protocol=Zero-shot2026.04 | 41.3 | 93.8 | |
| WandaBackbone=VideoLLaMA2 (7B), Sparsity Ratio=60%, Evaluation Protocol=Zero-shot2026.04 | 39.4 | 91 | |
| OWLBackbone=VideoLLaMA2 (7B), Sparsity Ratio=60%, Evaluation Protocol=Zero-shot2026.04 | 37.7 | 89.4 | |
| SparseGPTBackbone=VideoLLaMA2 (7B), Sparsity Ratio=60%, Evaluation Protocol=Zero-shot2026.04 | 35.7 | 88.5 | |
| VideoLLM-online-8BModel=VideoLLM-online-8B2026.03 | 26.9 | — | |
| MagnitudeSparsity Ratio=60%, Base Model=LLaVA-OneVision2025.04 | 20.7 | — | |
| MagnitudeBackbone=VideoLLaMA2 (7B), Sparsity Ratio=60%, Evaluation Protocol=Zero-shot2026.04 | 0 | 12.6 |