Vision-Language Understanding on MME
100Average ScoreVanilla
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| VanillaRetained Tokens=576, Token Reduction Ratio=100%2026.03 | 100 | 1,841 | — | — | |
| OriginalPrecision (Pre.)=FP16, Keep Ratio=100%, GFLOPs=100.00%2026.04 | 100 | 2,332.35 | — | — | |
| PRUNESIDRetained Tokens=192, Token Reduction Ratio=↓ 66.7%2026.03 | 98.7 | 1,842 | — | — | |
| VisionZipRetained Tokens=192, Token Reduction Ratio=↓ 66.7%2026.03 | 98.3 | 1,846 | — | — | |
| PruningPrecision (Pre.)=FP16, Keep Ratio=30%, GFLOPs=55.80%2026.04 | 97.68 | 2,278.27 | — | — | |
| PRUNESIDRetained Tokens=128, Token Reduction Ratio=↓ 77.8%2026.03 | 97.4 | 1,821 | — | — | |
| AWQPrecision (Pre.)=W4A16, Keep Ratio=100%, GFLOPs=100.00%2026.04 | 97 | — | — | — | |
| DivPrunePrecision (Pre.)=FP16, Keep Ratio=25%, GFLOPs=24.45%2026.04 | 96.95 | — | — | — | |
| GPTQPrecision (Pre.)=W4A16, Keep Ratio=100%, GFLOPs=100.00%2026.04 | 96.27 | — | — | — | |
| VisionZipRetained Tokens=128, Token Reduction Ratio=↓ 77.8%2026.03 | 96.2 | 1,841 | — | — | |
| LUQPrecision (Pre.)=W2.75A16, Keep Ratio=100%, GFLOPs=100.00%2026.04 | 95.93 | — | — | — | |
| Quant.Precision (Pre.)=W4A4, Keep Ratio=100%, GFLOPs=100.00%2026.04 | 95.9 | 2,236.61 | — | — | |
| QUOTAPrecision (Pre.)=W4A4, Keep Ratio=30%, GFLOPs=55.80%2026.04 | 94.52 | 2,204.52 | — | — | |
| PRUNESIDRetained Tokens=64, Token Reduction Ratio=↓ 88.9%2026.03 | 94.4 | 1,735 | — | — | |
| FastVPrecision (Pre.)=FP16, Keep Ratio=30%, GFLOPs=32.66%2026.04 | 94.36 | — | — | — | |
| VisionZipRetained Tokens=64, Token Reduction Ratio=↓ 88.9%2026.03 | 92.4 | 1,737 | — | — | |
| PDropPrecision (Pre.)=FP16, Keep Ratio=30%, GFLOPs=33.03%2026.04 | 86.5 | — | — | — | |
| VTWPrecision (Pre.)=FP16, Keep Ratio=40%, GFLOPs=43.43%2026.04 | 66.82 | — | — | — | |
| SMoE (k=2)k=2, Smoothing mechanism=None, Backbone=ViT-based, Model parameters=5.6B, Fine-tuning dataset=50% of the LLaVA-665K dataset, Number of experts=42026.06 | — | — | 40.18 | 67.78 | |
| SmoothSMoE (k=2.5)k=2.5, Smoothing mechanism=SmoothSMoE, Backbone=ViT-based, Model parameters=5.6B, Fine-tuning dataset=50% of the LLaVA-665K dataset, Number of experts=42026.06 | — | — | 40.85 | 70.93 | |
| SmoothSMoE annealed (k=2)k=2, Smoothing mechanism=Annealed SmoothSMoE, Backbone=ViT-based, Model parameters=5.6B, Fine-tuning dataset=50% of the LLaVA-665K dataset, Number of experts=42026.06 | — | — | 41.16 | 69.25 |