Vision-Language Performance on Average
80.46AccuracyQwen2.5-VL-7B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen2.5-VL-7BResolution=Full-resolution baseline2026.03 | 80.46 | 1 | |
| AwaResStrategy=Adaptive resolution escalation2026.03 | 80.3 | 0.36 | |
| VisionThinkStrategy=Adaptive resolution escalation2026.03 | 79.23 | 0.61 | |
| VisionZIPRetain Token Ratio Budget=70%2026.03 | 76.47 | 0.7 | |
| VisionZIPRetain Token Ratio Budget=50%2026.03 | 75.57 | 0.5 | |
| SparseVLMRetain Token Ratio Budget=70%2026.03 | 75.11 | 0.7 | |
| Holo-VRetain Token Ratio Budget=70%2026.03 | 73.92 | 0.7 | |
| SparseVLMRetain Token Ratio Budget=50%2026.03 | 73.46 | 0.5 | |
| Qwen2.5-VL-7BResolution=Low-res images2026.03 | 73.39 | 0.25 | |
| Holo-VRetain Token Ratio Budget=50%2026.03 | 69.86 | 0.5 |