Visual Question Answering on TextVQA (TextVQA Score)
90TextVQA ScoreSPOT-E
Evaluation Results
| Method | Links | |
|---|---|---|
| SPOT-EBase Model=Qwen3-VL-32B2026.06 | 90 | |
| SPOT-EBase Model=Qwen2.5-VL-32B2026.06 | 89.5 | |
| Qwen3-VL-32BBase Model=Qwen3-VL-32B2026.06 | 89 | |
| Qwen2.5-VL-32BBase Model=Qwen2.5-VL-32B2026.06 | 88.3 | |
| SPOT-EBase Model=LLaVA-OV-72B2026.06 | 88 | |
| LLaVA-OV-72BBase Model=LLaVA-OV-72B2026.06 | 86.4 | |
| SPOT-EBase Model=InternVL2.5-26B2026.06 | 86.1 | |
| InternVL2.5-26BBase Model=InternVL2.5-26B2026.06 | 84.6 | |
| Full PrecisionModel=Qwen2.5-VL-7B-Instruct, Weight bits=16 bit, Memory=15.45 GB, Evaluation protocol=zero-shot2026.07 | 83.88 | |
| Qwen3-VL-8BRetained Tokens=1024 Tokens (100%)2026.07 | 83.8 | |
| Full PrecisionModel=Qwen2.5-VL-72B-Instruct, Weight bits=16 bit, Memory=136.74 GB, Evaluation protocol=zero-shot2026.07 | 83.26 | |
| GPTQModel=Qwen2.5-VL-7B-Instruct, Weight bits=3 bit, Memory=2.66 GB, Evaluation protocol=zero-shot2026.07 | 82.25 | |
| GPTQModel=Qwen2.5-VL-72B-Instruct, Weight bits=3 bit, Memory=28.61 GB, Evaluation protocol=zero-shot2026.07 | 82.03 | |
| InternVL3-8BBackbone=InternVL3-8B, Retained Tokens=1280, Reduction Ratio=0%2026.06 | 81.5 | |
| SAB-LVLMModel=Qwen2.5-VL-72B-Instruct, Weight bits=1.07 bit, Memory=23.23 GB, Evaluation protocol=zero-shot2026.07 | 80.76 | |
| ARB-LLMModel=Qwen2.5-VL-72B-Instruct, Weight bits=1.19 bit, Memory=23.23 GB, Evaluation protocol=zero-shot2026.07 | 80.23 | |
| Full PrecisionModel=Qwen2.5-VL-32B-Instruct, Weight bits=16 bit, Memory=62.31 GB, Evaluation protocol=zero-shot2026.07 | 78.85 | |
| BiLLMModel=Qwen2.5-VL-72B-Instruct, Weight bits=1.08 bit, Memory=24.23 GB, Evaluation protocol=zero-shot2026.07 | 77.28 | |
| SAB-LVLMModel=Qwen2.5-VL-32B-Instruct, Weight bits=1.07 bit, Memory=10.28 GB, Evaluation protocol=zero-shot2026.07 | 75.45 | |
| EADPRetained Tokens=512 Tokens (↓ 50.0%)2026.07 | 75.2 | |
| DivPruneRetained Tokens=512 Tokens (↓ 50.0%)2026.07 | 75 | |
| ARB-LLMModel=Qwen2.5-VL-32B-Instruct, Weight bits=1.16 bit, Memory=10.28 GB, Evaluation protocol=zero-shot2026.07 | 74.96 | |
| CDPrunerRetained Tokens=512 Tokens (↓ 50.0%)2026.07 | 74.5 | |
| GPTQModel=Qwen2.5-VL-32B-Instruct, Weight bits=3 bit, Memory=12.71 GB, Evaluation protocol=zero-shot2026.07 | 74.45 | |
| PB-LLMModel=Qwen2.5-VL-72B-Instruct, Weight bits=1.70 bit, Memory=22.99 GB, Evaluation protocol=zero-shot2026.07 | 74.34 | |
| SAB-LVLMModel=Qwen2.5-VL-7B-Instruct, Weight bits=1.07 bit, Memory=2.14 GB, Evaluation protocol=zero-shot2026.07 | 74 | |
| EADPRetained Tokens=256 Tokens (↓ 75.0%)2026.07 | 71.4 | |
| ERABackbone=InternVL3-8B, Retained Tokens=256, Reduction Ratio=80.0%2026.06 | 71.2 | |
| ARB-LLMModel=Qwen2.5-VL-7B-Instruct, Weight bits=1.07 bit, Memory=2.14 GB, Evaluation protocol=zero-shot2026.07 | 70.7 | |
| DivPruneRetained Tokens=256 Tokens (↓ 75.0%)2026.07 | 69.5 | |
| CDPrunerRetained Tokens=256 Tokens (↓ 75.0%)2026.07 | 68.5 | |
| VisionZipBackbone=InternVL3-8B, Retained Tokens=256, Reduction Ratio=80.0%2026.06 | 67.2 | |
| DivPruneBackbone=InternVL3-8B, Retained Tokens=256, Reduction Ratio=80.0%2026.06 | 67.1 | |
| EADPRetained Tokens=128 Tokens (↓ 87.5%)2026.07 | 66.1 | |
| PB-LLMModel=Qwen2.5-VL-32B-Instruct, Weight bits=1.70 bit, Memory=10.24 GB, Evaluation protocol=zero-shot2026.07 | 63.62 | |
| DivPruneRetained Tokens=128 Tokens (↓ 87.5%)2026.07 | 62.6 | |
| DARTBackbone=InternVL3-8B, Retained Tokens=256, Reduction Ratio=80.0%2026.06 | 62 | |
| CleanModel=LLaVA-1.5, Vision Encoder=CLIP-3362024.12 | 61.3 | |
| ERABackbone=InternVL3-8B, Retained Tokens=128, Reduction Ratio=90.0%2026.06 | 60.8 | |
| CDPrunerRetained Tokens=128 Tokens (↓ 87.5%)2026.07 | 60.6 | |
| CleanModel=Qwen2.5-VL, Vision Encoder=Qwen-ViT2024.12 | 59.9 | |
| Full-DataSampling Ratio=100% (665K), Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 57.9 | |
| Both-EmbModel=LLaVA-1.5, Vision Encoder=CLIP-3362024.12 | 56.2 | |
| CoIDOSampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 56 | |
| OFASampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 56 | |
| BiLLMModel=Qwen2.5-VL-32B-Instruct, Weight bits=1.08 bit, Memory=10.73 GB, Evaluation protocol=zero-shot2026.07 | 55.92 | |
| Img-EmbModel=LLaVA-1.5, Vision Encoder=CLIP-3362024.12 | 55.4 | |
| PreSelSampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 55.2 | |
| XMASSampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 55.2 | |
| ICONSSampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 55.2 | |
| COINCIDESampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 54.8 | |
| DivPruneBackbone=InternVL3-8B, Retained Tokens=128, Reduction Ratio=90.0%2026.06 | 54.7 | |
| RandomSampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 54.3 | |
| Text-EmbModel=LLaVA-1.5, Vision Encoder=CLIP-3362024.12 | 54.1 | |
| CLIP-ScoreSampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 53.4 | |
| BiLLMModel=Qwen2.5-VL-7B-Instruct, Weight bits=1.08 bit, Memory=2.24 GB, Evaluation protocol=zero-shot2026.07 | 53.38 | |
| TypiClustSampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 53.3 | |
| IFDSampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 51.8 | |
| Self-FilterSampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 51.4 | |
| DARTBackbone=InternVL3-8B, Retained Tokens=128, Reduction Ratio=90.0%2026.06 | 51.3 | |
| EL2NSampling Ratio=15%, Backbone=LLaVA-v1.5-7B, Training Dataset=LLaVA-665K2026.05 | 50.2 | |
| CleanModel=LLaVA, Vision Encoder=CLIP-3362024.12 | 50.2 | |
| VisionZipBackbone=InternVL3-8B, Retained Tokens=128, Reduction Ratio=90.0%2026.06 | 48.3 | |
| Img-EmbModel=LLaVA, Vision Encoder=CLIP-3362024.12 | 48 | |
| Img-EmbModel=Qwen2.5-VL, Vision Encoder=Qwen-ViT2024.12 | 47.7 | |
| Both-EmbModel=LLaVA, Vision Encoder=CLIP-3362024.12 | 47.4 | |
| Text-EmbModel=LLaVA, Vision Encoder=CLIP-3362024.12 | 44.8 | |
| VEV-UAPModel=Qwen2.5-VL, Vision Encoder=Qwen-ViT2024.12 | 43.8 | |
| CleanModel=InstructBlip, Vision Encoder=EVA-CLIP2024.12 | 32.5 | |
| Img-EmbModel=InstructBlip, Vision Encoder=EVA-CLIP2024.12 | 30.1 | |
| Both-EmbModel=InstructBlip, Vision Encoder=EVA-CLIP2024.12 | 29.6 | |
| PB-LLMModel=Qwen2.5-VL-7B-Instruct, Weight bits=1.17 bit, Memory=2.15 GB, Evaluation protocol=zero-shot2026.07 | 28.67 | |
| Text-EmbModel=InstructBlip, Vision Encoder=EVA-CLIP2024.12 | 26.6 | |
| VEV-UAPModel=InstructBlip, Vision Encoder=EVA-CLIP2024.12 | 26.3 | |
| VEV-UAPModel=LLaVA, Vision Encoder=CLIP-3362024.12 | 2.9 | |
| VEV-UAPModel=LLaVA-1.5, Vision Encoder=CLIP-3362024.12 | 0.6 |