Multi-modal Understanding on MuirBench
59.6ScoreVanilla
Evaluation Results
| Method | Links | |
|---|---|---|
| VanillaMode=Upper bound2024.12 | 59.6 | |
| iLLaVAImage token reduction ratio=66.7%2024.12 | 59.1 | |
| PyramidDropImage token reduction ratio=66.7%2024.12 | 58.7 | |
| VisionZipImage token reduction ratio=66.7%2024.12 | 58.2 | |
| SparseVLMImage token reduction ratio=66.7%2024.12 | 58.1 | |
| FasterVLMImage token reduction ratio=66.7%2024.12 | 57.8 | |
| VisionZipImage token reduction ratio=77.8%2024.12 | 57.6 | |
| iLLaVAImage token reduction ratio=77.8%2024.12 | 57.2 | |
| SparseVLMImage token reduction ratio=77.8%2024.12 | 56.4 | |
| FasterVLMImage token reduction ratio=77.8%2024.12 | 56.3 | |
| iLLaVAImage token reduction ratio=88.9%2024.12 | 56.3 | |
| PyramidDropImage token reduction ratio=77.8%2024.12 | 55.3 | |
| FasterVLMImage token reduction ratio=88.9%2024.12 | 55.2 | |
| VisionZipImage token reduction ratio=88.9%2024.12 | 54.8 | |
| SparseVLMImage token reduction ratio=88.9%2024.12 | 54.6 | |
| PyramidDropImage token reduction ratio=88.9%2024.12 | 53 |