Optical Character Recognition on OCRBench (accuracy)
498AccuracyRADIOv2.5-H
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| RADIOv2.5-HResolution=768^2, Tokens/im=196, LLM=MN-Minitron-8B, Data mixture=LLaVA1.62024.12 | 498 | — | — | |
| RADIOv2.5-gResolution=672^2, Tokens/im=196, LLM=MN-Minitron-8B, Data mixture=LLaVA1.62024.12 | 479 | — | — | |
| SigLIP-SO400MResolution=384^2, Tokens/im=196, LLM=MN-Minitron-8B, Data mixture=LLaVA1.62024.12 | 452 | — | — | |
| RADIOv2.5-LResolution=768^2, Tokens/im=196, LLM=MN-Minitron-8B, Data mixture=LLaVA1.62024.12 | 441 | — | — | |
| OpenAI-CLIPResolution=336^2, Tokens/im=196, LLM=MN-Minitron-8B, Data mixture=LLaVA1.62024.12 | 431 | — | — | |
| RADIOv2.1-HResolution=512^2, Tokens/im=196, LLM=MN-Minitron-8B, Data mixture=LLaVA1.62024.12 | 392 | — | — | |
| RADIOv2.1-HResolution=768^2, Tokens/im=196, LLM=MN-Minitron-8B, Data mixture=LLaVA1.62024.12 | 365 | — | — | |
| Qwen2.5-VL-7B + ECRDParameters=7B, Enhancement=ECRD2026.02 | 90.7 | — | — | |
| Qwen3-VL-8B-InstructContext Window=32K2026.03 | 90 | — | — | |
| Qwen3-VL-8B-InstructContext Window=4K2026.03 | 89.2 | — | — | |
| Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B, Token Retention Ratio=100%2025.12 | 88.5 | — | — | |
| Qwen3-VL-32B-InstructContext Window=4K2026.03 | 88.5 | — | — | |
| Qwen3-VL-32B-InstructContext Window=32K2026.03 | 88.5 | — | — | |
| VanillaAvg. Tokens=100%2026.02 | 88.4 | — | — | |
| Kimi-VL-A3B-Instruct2026.03 | 86.5 | — | — | |
| Qwen2.5-VL-7B + supervisorParameters=7B, Enhancement=supervisor2026.02 | 85.7 | — | — | |
| BaselineCompression Ratio=0%, Backbone=Qwen-2.5-7B-Instruct2026.02 | 85.3 | — | — | |
| Qwen2-VL-72B-ThinkingMax Output Tokens=40K2026.03 | 85 | — | — | |
| RADIO1DTokens per frame/tile=224, TTFT (ms)=410.2, LLM Backbone=9B Nemotron2026.07 | 85 | — | — | |
| RADIO1DTokens per frame/tile=256, TTFT (ms)=452.8, LLM Backbone=9B Nemotron2026.07 | 84.4 | — | — | |
| C-RADIOv4-HTokens per frame/tile=256, TTFT (ms)=468.2, LLM Backbone=9B Nemotron2026.07 | 84.3 | — | — | |
| DivPruneBackbone=Qwen2.5-VL-7B, Token Retention Ratio=50%2025.12 | 84.1 | — | — | |
| Qwen2-VL-72B-ThinkingMax Output Tokens=4K2026.03 | 83.7 | — | — | |
| RADIO1DTokens per frame/tile=192, TTFT (ms)=380.4, LLM Backbone=9B Nemotron2026.07 | 83.5 | — | — | |
| DivPrune + RandomBackbone=Qwen2.5-VL-7B, Token Retention Ratio=50%2025.12 | 83.3 | — | — | |
| RADIO1DTokens per frame/tile=128, TTFT (ms)=373.2, LLM Backbone=9B Nemotron2026.07 | 82.8 | — | — | |
| SigLIP2-SO400mTokens per frame/tile=256, TTFT (ms)=440.5, LLM Backbone=9B Nemotron2026.07 | 82.7 | — | — | |
| SigLIP2-gTokens per frame/tile=256, TTFT (ms)=517.6, LLM Backbone=9B Nemotron2026.07 | 82.5 | — | — | |
| Qwen2.5-VL-7BParameters=7B2026.02 | 82.3 | — | — | |
| Qwen2-VL-7B-ThinkingMax Output Tokens=40K2026.03 | 82 | — | — | |
| Qwen2-VL-7B-ThinkingMax Output Tokens=4K2026.03 | 81.2 | — | — | |
| RADIO1DTokens per frame/tile=64, TTFT (ms)=346.1, LLM Backbone=9B Nemotron2026.07 | 81.2 | — | — | |
| Baseline2026.02 | 80.9 | — | — | |
| fp16Model=Qwen2-VL-7b, Bitwidth=fp162026.02 | 80.7 | — | — | |
| DivPrune+VTWBackbone=Qwen2.5-VL-7B, Token Retention Ratio=50%2025.12 | 79.9 | — | — | |
| Kimi-VL-A3B-Thinking2026.03 | 79.9 | — | — | |
| DART + RandomBackbone=Qwen2.5-VL-7B, Token Retention Ratio=50%2025.12 | 77.9 | — | — | |
| RADIO1DTokens per frame/tile=32, TTFT (ms)=335.1, LLM Backbone=9B Nemotron2026.07 | 77.8 | — | — | |
| fp16Model=Qwen2-VL-2b, Bitwidth=fp162026.02 | 76.5 | — | — | |
| Phi-4-reasoning-vision-15B2026.03 | 76 | — | — | |
| Phi-4-reasoning-vision-15BThinking Mode=default2026.03 | 76 | — | — | |
| Phi-4-reasoning-vision-15BProtocol=force nothink2026.03 | 75.6 | — | — | |
| DARTBackbone=Qwen2.5-VL-7B, Token Retention Ratio=50%2025.12 | 75.5 | — | — | |
| gemma-3-12b-it2026.03 | 75.3 | — | — | |
| gemma-2-9b-it2026.03 | 75.3 | — | — | |
| TLQModel=Qwen2-VL-7b, Bitwidth=W4A82026.02 | 74.2 | — | — | |
| RADIO1DTokens per frame/tile=8, TTFT (ms)=329.7, LLM Backbone=9B Nemotron2026.07 | 74.2 | — | — | |
| IDPrunerCompression Ratio=75%, Backbone=Qwen-2.5-7B-Instruct2026.02 | 74 | — | — | |
| LLaVA-OneVision-7B + ECRDParameters=7B, Enhancement=ECRD2026.02 | 73.8 | — | — | |
| TTAugstrategy=Test-Time Augmentation2025.10 | 73.7 | — | — | |
| TTAugAdaptation strategy=Test-time Augmentation2025.10 | 73.7 | — | — | |
| Phi-4-reasoning-vision-15BThinking Mode=force thinking2026.03 | 73.7 | — | — | |
| DART+VTWBackbone=Qwen2.5-VL-7B, Token Retention Ratio=50%2025.12 | 73.3 | — | — | |
| TTAdaptstrategy=Model parameter adaptation (1)2025.10 | 73 | — | — | |
| (1)Adaptation strategy=(1)2025.10 | 73 | — | — | |
| Baselinestrategy=Baseline2025.10 | 72.9 | — | — | |
| Baseline2025.10 | 72.9 | — | — | |
| MBQModel=Qwen2-VL-7b, Bitwidth=W4A82026.02 | 72.8 | — | — | |
| VisionSelectorCompression Ratio=75%, Backbone=Qwen-2.5-7B-Instruct2026.02 | 72.5 | — | — | |
| TTAdaptstrategy=Model parameter adaptation (2)2025.10 | 70.5 | — | — | |
| (2)Adaptation strategy=Model parameter adaptation2025.10 | 70.5 | — | — | |
| TLQModel=Qwen2-VL-7b, Bitwidth=W4A62026.02 | 68.4 | — | — | |
| FastVAvg. Tokens=52%2026.02 | 67.3 | — | — | |
| TLQModel=Qwen2-VL-2b, Bitwidth=W4A82026.02 | 66.7 | — | — | |
| IVC-PruneAvg. Tokens=50%2026.02 | 66.3 | — | — | |
| RTNModel=Qwen2-VL-2b, Bitwidth=W4A82026.02 | 66.2 | — | — | |
| MBQModel=Qwen2-VL-7b, Bitwidth=W4A62026.02 | 65.8 | — | — | |
| LLaVA next 8B# Vis tok.=28802024.12 | 65.4 | — | — | |
| MBQModel=Qwen2-VL-2b, Bitwidth=W4A82026.02 | 65.4 | — | — | |
| TLQModel=Qwen2-VL-2b, Bitwidth=W4A62026.02 | 63.5 | — | — | |
| Florence-VL 8B# Vis tok.=5762024.12 | 63.4 | — | — | |
| Florence-VL 3B# Vis tok.=5762024.12 | 63 | — | — | |
| Phi-4-mm-instruct2026.03 | 62.6 | — | — | |
| Cambrian 8B# Vis tok.=5762024.12 | 62.4 | — | — | |
| fp16Model=LLaVA-onevision-7b, Bitwidth=fp162026.02 | 62.2 | — | — | |
| LLaVA-OneVision-7BParameters=7B2026.02 | 62.2 | — | — | |
| RTNModel=Qwen2-VL-7b, Bitwidth=W4A62026.02 | 62 | — | — | |
| PDropAvg. Tokens=52%2026.02 | 62 | — | — | |
| SCOPECompression Ratio=75%, Backbone=Qwen-2.5-7B-Instruct2026.02 | 61.7 | — | — | |
| RTNModel=Qwen2-VL-2b, Bitwidth=W4A62026.02 | 61 | — | — | |
| RTNModel=Qwen2-VL-7b, Bitwidth=W4A82026.02 | 60.3 | — | — | |
| Phi 3.5 Vision2024.12 | 59.8 | — | — | |
| MBQModel=Qwen2-VL-2b, Bitwidth=W4A62026.02 | 59.8 | — | — | |
| SmoothQuantModel=Qwen2-VL-2b, Bitwidth=W4A62026.02 | 59.4 | — | — | |
| DivPruneCompression Ratio=75%, Backbone=Qwen-2.5-7B-Instruct2026.02 | 58.4 | — | — | |
| IDPrunerRetain Tokens=25%2026.02 | 57.8 | — | — | |
| fp16Model=LLaVA-onevision-0.5b, Bitwidth=fp162026.02 | 57.6 | — | — | |
| SmoothQuantModel=Qwen2-VL-2b, Bitwidth=W4A82026.02 | 57.6 | — | — | |
| FusedModel size=7B, Evaluation protocol=0-shot2026.02 | 57.6 | — | — | |
| MaD-MixModel size=7B, Evaluation protocol=0-shot2026.02 | 57.2 | — | — | |
| SmoothQuantModel=Qwen2-VL-7b, Bitwidth=W4A82026.02 | 57.1 | — | — | |
| AvgModel size=7B, Evaluation protocol=0-shot2026.02 | 56.9 | — | — | |
| UniformModel size=7B, Evaluation protocol=0-shot2026.02 | 56.5 | — | — | |
| RADIO1DTokens per frame/tile=1, TTFT (ms)=327.0, LLM Backbone=9B Nemotron2026.07 | 56.5 | — | — | |
| SmoothQuantModel=Qwen2-VL-7b, Bitwidth=W4A62026.02 | 55.9 | — | — | |
| VisionSelectorCompression Ratio=90%, Backbone=Qwen-2.5-7B-Instruct2026.02 | 55.5 | — | — | |
| VisionSelectorRetain Tokens=25%2026.02 | 55.5 | — | — | |
| IDPrunerCompression Ratio=90%, Backbone=Qwen-2.5-7B-Instruct2026.02 | 53.9 | — | — | |
| TLQModel=LLaVA-onevision-7b, Bitwidth=W4A82026.02 | 53.3 | — | — | |
| MBQModel=LLaVA-onevision-7b, Bitwidth=W4A82026.02 | 52.3 | — | — |