Text-focused Image Understanding on Text-focused Image Understanding Benchmarks (test)
84.1Chart ScoreInternVL2.5-8B
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| InternVL2.5-8BToken Reduction Ratio=0% (Upper Bound), Backbone=InternVL2.5-8B2026.04 | 84.1 | 80.3 | 92 | 69 | 100 | |
| Qwen2.5-VL-7BToken Reduction Ratio=0% (Upper Bound), Backbone=Qwen2.5-VL-7B2026.04 | 84 | 83.8 | 95 | 81 | 100 | |
| ID-SelectionToken Reduction Ratio=77.8%, Backbone=InternVL2.5-8B, Importance Mechanism=LLM second-layer cross-modal attention2026.04 | 73.2 | 60.2 | 68 | 44 | 74.9 | |
| ID-SelectionToken Reduction Ratio=77.8%, Backbone=Qwen2.5-VL-7B, Importance Mechanism=LLM second-layer cross-modal attention2026.04 | 72.4 | 72.8 | 86 | 60 | 84.4 | |
| FastVToken Reduction Ratio=77.8%, Backbone=Qwen2.5-VL-7B2026.04 | 65.4 | 58.7 | 79 | 54 | 74.4 | |
| FastVToken Reduction Ratio=77.8%, Backbone=InternVL2.5-8B2026.04 | 65.2 | 56.4 | 59 | 40 | 67.5 | |
| DARTToken Reduction Ratio=77.8%, Backbone=InternVL2.5-8B2026.04 | 65 | 59.9 | 58 | 39 | 67.8 | |
| ID-SelectionToken Reduction Ratio=88.9%, Backbone=InternVL2.5-8B, Importance Mechanism=LLM second-layer cross-modal attention2026.04 | 64.2 | 49 | 51 | 34 | 60.5 | |
| ID-SelectionToken Reduction Ratio=88.9%, Backbone=Qwen2.5-VL-7B, Importance Mechanism=LLM second-layer cross-modal attention2026.04 | 60.7 | 62 | 76 | 45 | 70.5 | |
| DARTToken Reduction Ratio=77.8%, Backbone=Qwen2.5-VL-7B2026.04 | 53.2 | 58.8 | 60 | 40 | 61.5 | |
| FastVToken Reduction Ratio=88.9%, Backbone=Qwen2.5-VL-7B2026.04 | 43 | 42.8 | 60 | 39 | 53.4 | |
| DARTToken Reduction Ratio=88.9%, Backbone=InternVL2.5-8B2026.04 | 42.7 | 44.3 | 38 | 31 | 48 | |
| FastVToken Reduction Ratio=88.9%, Backbone=InternVL2.5-8B2026.04 | 38.2 | 36.2 | 38 | 31 | 44.2 | |
| DARTToken Reduction Ratio=88.9%, Backbone=Qwen2.5-VL-7B2026.04 | 37.4 | 44.2 | 41 | 31 | 44.7 |