Multi-modal Understanding on MMMU (Specific Accuracy Reduction Ratios)
64.3Accuracy (77.8% reduction ratio)iLLaVA
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| iLLaVAStyle=Token merging2024.12 | 64.3 | 62.8 | |
| AIMStyle=Token merging2024.12 | 63.8 | 61.2 | |
| VisionZipStyle=Token merging2024.12 | 63.6 | 61.4 | |
| FEATHERStyle=Token pruning2024.12 | 63.4 | 61 | |
| DivPruneStyle=Token pruning2024.12 | 63.1 | 60.8 | |
| FasterVLMStyle=Token pruning2024.12 | 62.9 | 61 | |
| AdaFVStyle=Token pruning2024.12 | 62.4 | 59.8 | |
| PyramidDropStyle=Token pruning2024.12 | 62.2 | 57.2 | |
| SparseVLMStyle=Token merging2024.12 | 62.2 | 60 | |
| LLaVA-PruMergeStyle=Token merging2024.12 | 61.8 | 59.6 | |
| FastVStyle=Token pruning2024.12 | 58.2 | 52.4 |