Multi-modal Understanding on MME (Score, Rel. Perf.)
1,842MME ScoreLLaVA-NeXT-7B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| LLaVA-NeXT-7BVisual tokens retained=Full, Pruning ratio=0%, Base Model=LLaVA-NeXT-7B2026.03 | 1,842 | 100 | |
| ResPruneVisual tokens retained=640, Pruning ratio=66.7%, Base Model=LLaVA-NeXT-7B2026.03 | 1,821 | 99.6 | |
| FastVVisual tokens retained=640, Pruning ratio=66.7%, Base Model=LLaVA-NeXT-7B2026.03 | 1,807 | 94.7 | |
| DARTVisual tokens retained=640, Pruning ratio=66.7%, Base Model=LLaVA-NeXT-7B2026.03 | 1,793 | 97.5 | |
| PruMergeVisual tokens retained=640, Pruning ratio=66.7%, Base Model=LLaVA-NeXT-7B2026.03 | 1,790 | 96.6 | |
| PDropVisual tokens retained=640, Pruning ratio=66.7%, Base Model=LLaVA-NeXT-7B2026.03 | 1,782 | 95.7 | |
| VisionZIPVisual tokens retained=640, Pruning ratio=66.7%, Base Model=LLaVA-NeXT-7B2026.03 | 1,782 | 98.1 | |
| ResPruneVisual tokens retained=320, Pruning ratio=88.9%, Base Model=LLaVA-NeXT-7B2026.03 | 1,780 | 98.1 | |
| DivPruneVisual tokens retained=640, Pruning ratio=66.7%, Base Model=LLaVA-NeXT-7B2026.03 | 1,773 | 97.2 | |
| SparseVLMVisual tokens retained=640, Pruning ratio=66.7%, Base Model=LLaVA-NeXT-7B2026.03 | 1,772 | 96.9 | |
| SparseVLMVisual tokens retained=320, Pruning ratio=88.9%, Base Model=LLaVA-NeXT-7B2026.03 | 1,747 | 93.1 | |
| PruMergeVisual tokens retained=320, Pruning ratio=88.9%, Base Model=LLaVA-NeXT-7B2026.03 | 1,744 | 94.1 | |
| DivPruneVisual tokens retained=320, Pruning ratio=88.9%, Base Model=LLaVA-NeXT-7B2026.03 | 1,731 | 96.6 | |
| DARTVisual tokens retained=320, Pruning ratio=88.9%, Base Model=LLaVA-NeXT-7B2026.03 | 1,710 | 94.8 | |
| VisionZIPVisual tokens retained=320, Pruning ratio=88.9%, Base Model=LLaVA-NeXT-7B2026.03 | 1,698 | 95 | |
| PDropVisual tokens retained=320, Pruning ratio=88.9%, Base Model=LLaVA-NeXT-7B2026.03 | 1,672 | 83.5 | |
| FastVVisual tokens retained=320, Pruning ratio=88.9%, Base Model=LLaVA-NeXT-7B2026.03 | 1,539 | 80.7 |