OCR-based Visual Question Answering on OCR-VQA
65.6AccuracyCapPa L/14
Evaluation Results
| Method | Links | |
|---|---|---|
| CapPa L/14Backbone Scale=L/14, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 65.6 | |
| CLIP L/14Backbone Scale=L/14, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 64.1 | |
| CapBackbone Scale=B/16, Pre-training Batch Size=8k, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 62.2 | |
| CapPaBackbone Scale=B/16, Pre-training Batch Size=8k, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 62.2 | |
| CLIP* L/14Backbone Scale=L/14, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 61.3 | |
| CLIPBackbone Scale=B/16, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 60 | |
| CLIP* (16k)Backbone Scale=B/16, Pre-training Batch Size=16k, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 56.5 | |
| CLIP* (8k)Backbone Scale=B/16, Pre-training Batch Size=8k, Evaluation Protocol=Frozen visual representations with single transformer decoder2023.06 | 56 | |
| VanillaBackbone=LLaVA-Next-7B, Token Reduction Rate=0.0%2025.05 | 51.7 | |
| TrimTokenator-LCBackbone=LLaVA-1.5-13B, Retention Ratio=0.22025.12 | 49.5 | |
| SparseVLMBackbone=LLaVA-1.5-13B, Retention Ratio=0.22025.12 | 47.5 | |
| VisionZipBackbone=LLaVA-1.5-13B, Retention Ratio=0.22025.12 | 47 | |
| TrimTokenatorBackbone=LLaVA-1.5-13B, Retention Ratio=0.22025.12 | 47 | |
| DivPruneBackbone=LLaVA-1.5-13B, Retention Ratio=0.22025.12 | 46.5 | |
| Original ModelBackbone=LLaVA-1.5-13B, Retention Ratio=1.02025.12 | 46 | |
| VTWBackbone=LLaVA-1.5-13B, Retention Ratio=0.22025.12 | 46 | |
| FastVBackbone=LLaVA-1.5-13B, Retention Ratio=0.22025.12 | 45.5 | |
| CATPBackbone=LLaVA-1.5-13B, Retention Ratio=0.22025.12 | 45.5 | |
| MoB (with η-prior)Backbone=LLaVA-Next-7B, Objectives=PA VP, Pruning budget K=320, Token Reduction Rate=88.9%2025.05 | 41.8 | |
| DARTBackbone=LLaVA-Next-7B, Objectives=VP, Pruning budget K=320, Token Reduction Rate=88.9%2025.05 | 40.6 | |
| FasterVLMBackbone=LLaVA-Next-7B, Objectives=VP, Pruning budget K=320, Token Reduction Rate=88.9%2025.05 | 40.1 | |
| MustDropBackbone=LLaVA-Next-7B, Objectives=PA VP, Pruning budget K=320, Token Reduction Rate=88.9%2025.05 | 38.2 | |
| FastVBackbone=LLaVA-Next-7B, Objectives=VP, Pruning budget K=320, Token Reduction Rate=88.9%2025.05 | 37.4 | |
| MoB (+ η-prior)Backbone=LLaVA-1.5-7B, Objectives=-, Pruning budget K=192, Token Reduction Rate=66.7%2025.05 | 30.7 | |
| MoB (w/o η-prior)Backbone=LLaVA-1.5-7B, Objectives=PA VP, Pruning budget K=192, Token Reduction Rate=66.7%2025.05 | 30.4 | |
| MoB (w/o η-prior)Backbone=LLaVA-1.5-7B, Objectives=PA VP, Pruning budget K=128, Token Reduction Rate=77.8%2025.05 | 29.9 | |
| MoB (+ η-prior)Backbone=LLaVA-1.5-7B, Objectives=-, Pruning budget K=128, Token Reduction Rate=77.8%2025.05 | 29.9 | |
| VanillaBackbone=LLaVA-1.5-7B, Token Reduction Rate=0.0%2025.05 | 29.7 | |
| DARTBackbone=LLaVA-1.5-7B, Objectives=VP, Pruning budget K=192, Token Reduction Rate=66.7%2025.05 | 29.6 | |
| DARTBackbone=LLaVA-1.5-7B, Objectives=VP, Pruning budget K=128, Token Reduction Rate=77.8%2025.05 | 29.6 | |
| SparseVLMBackbone=LLaVA-1.5-7B, Objectives=PA, Pruning budget K=192, Token Reduction Rate=66.7%2025.05 | 29.2 | |
| FastVBackbone=LLaVA-1.5-7B, Objectives=VP, Pruning budget K=192, Token Reduction Rate=66.7%2025.05 | 29.1 | |
| MustDropBackbone=LLaVA-1.5-7B, Objectives=PA VP, Pruning budget K=192, Token Reduction Rate=66.7%2025.05 | 28.9 | |
| FastVBackbone=LLaVA-1.5-7B, Objectives=VP, Pruning budget K=128, Token Reduction Rate=77.8%2025.05 | 28.5 | |
| MustDropBackbone=LLaVA-1.5-7B, Objectives=PA VP, Pruning budget K=128, Token Reduction Rate=77.8%2025.05 | 28.1 | |
| SparseVLMBackbone=LLaVA-1.5-7B, Objectives=PA, Pruning budget K=128, Token Reduction Rate=77.8%2025.05 | 28 | |
| MoB (w/o η-prior)Backbone=LLaVA-1.5-7B, Objectives=PA VP, Pruning budget K=64, Token Reduction Rate=88.9%2025.05 | 27.7 | |
| MoB (+ η-prior)Backbone=LLaVA-1.5-7B, Objectives=-, Pruning budget K=64, Token Reduction Rate=88.9%2025.05 | 27.7 | |
| Original ModelBackbone=Yi-VL-6B, Retention Ratio=1.02025.12 | 27.5 | |
| DARTBackbone=LLaVA-1.5-7B, Objectives=VP, Pruning budget K=64, Token Reduction Rate=88.9%2025.05 | 27 | |
| SparseVLMBackbone=LLaVA-Next-7B, Objectives=PA, Pruning budget K=320, Token Reduction Rate=88.9%2025.05 | 27 | |
| MustDropBackbone=LLaVA-1.5-7B, Objectives=PA VP, Pruning budget K=64, Token Reduction Rate=88.9%2025.05 | 26.7 | |
| TrimTokenator-LCBackbone=Yi-VL-6B, Retention Ratio=0.22025.12 | 25.5 | |
| DivPruneBackbone=Yi-VL-6B, Retention Ratio=0.22025.12 | 25 | |
| SparseVLMBackbone=Yi-VL-6B, Retention Ratio=0.22025.12 | 24.5 | |
| FastVBackbone=LLaVA-1.5-7B, Objectives=VP, Pruning budget K=64, Token Reduction Rate=88.9%2025.05 | 24.5 | |
| VisionZipBackbone=Yi-VL-6B, Retention Ratio=0.22025.12 | 24 | |
| TrimTokenatorBackbone=Yi-VL-6B, Retention Ratio=0.22025.12 | 24 | |
| VTWBackbone=Yi-VL-6B, Retention Ratio=0.22025.12 | 23.5 | |
| CATPBackbone=Yi-VL-6B, Retention Ratio=0.22025.12 | 23.5 | |
| FastVBackbone=Yi-VL-6B, Retention Ratio=0.22025.12 | 23 | |
| SparseVLMBackbone=LLaVA-1.5-7B, Objectives=PA, Pruning budget K=64, Token Reduction Rate=88.9%2025.05 | 18 | |
| Original ModelBackbone=LLaVA-1.5-7B, Retention Ratio=1.02025.12 | 8 | |
| TrimTokenator-LCBackbone=LLaVA-1.5-7B, Retention Ratio=0.22025.12 | 4.5 | |
| FastVBackbone=LLaVA-1.5-7B, Retention Ratio=0.22025.12 | 3.5 | |
| CATPBackbone=LLaVA-1.5-7B, Retention Ratio=0.22025.12 | 3.5 | |
| TrimTokenatorBackbone=LLaVA-1.5-7B, Retention Ratio=0.22025.12 | 3.5 | |
| SparseVLMBackbone=LLaVA-1.5-7B, Retention Ratio=0.22025.12 | 3 | |
| VisionZipBackbone=LLaVA-1.5-7B, Retention Ratio=0.22025.12 | 3 | |
| DivPruneBackbone=LLaVA-1.5-7B, Retention Ratio=0.22025.12 | 3 | |
| VTWBackbone=LLaVA-1.5-7B, Retention Ratio=0.22025.12 | 2.5 |