Vision-Centric Reasoning on V* Bench (Overall)
96.5Attribute ScoreQwen3-VL-4B + SD-RPN
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen3-VL-4B + SD-RPNRelative inference throughput (Tp)=0.58×, Max visual tokens=4,0962026.04 | 96.5 | 82.9 | 91.1 | |
| ZwZ-Qwen3-VL-4B + Q-ZoomRelative inference throughput (Tp)=0.66×, Max visual tokens=4,0962026.04 | 96.5 | 90.8 | 94.2 | |
| Qwen3-VL-4B + Q-ZoomRelative inference throughput (Tp)=0.73×, Max visual tokens=4,0962026.04 | 95.7 | 85.5 | 91.6 | |
| Qwen2.5-VL-7B + SD-RPNRelative inference throughput (Tp)=0.77×, Max visual tokens=4,0962026.04 | 94.8 | 77.6 | 88 | |
| ZwZ-Qwen2.5-VL-7B + Q-ZoomRelative inference throughput (Tp)=0.76×, Max visual tokens=4,0962026.04 | 94.8 | 86.8 | 91.6 | |
| Qwen2.5-VL-3B + SD-RPNRelative inference throughput (Tp)=0.66×, Max visual tokens=4,0962026.04 | 91.3 | 65.8 | 81.2 | |
| Qwen2.5-VL-7B + Q-ZoomRelative inference throughput (Tp)=0.86×, Max visual tokens=4,0962026.04 | 89.6 | 79 | 85.3 | |
| ZwZ-Qwen3-VL-4BRelative inference throughput (Tp)=1.00×, Max visual tokens=4,0962026.04 | 89.6 | 86.8 | 88.5 | |
| Qwen2.5-VL-3B + Q-ZoomRelative inference throughput (Tp)=0.67×, Max visual tokens=4,0962026.04 | 87 | 69.7 | 80.1 | |
| ZwZ-Qwen2.5-VL-7BRelative inference throughput (Tp)=1.00×, Max visual tokens=4,0962026.04 | 87 | 82.9 | 85.3 | |
| Qwen3-VL-4BRelative inference throughput (Tp)=1.00×, Max visual tokens=4,0962026.04 | 86.1 | 76.3 | 83.2 | |
| Qwen2.5-VL-7B + ThymeRelative inference throughput (Tp)=0.21×, Max visual tokens=4,0962026.04 | 83.5 | 80.3 | 82.2 | |
| Qwen2.5-VL-3BRelative inference throughput (Tp)=1.00×, Max visual tokens=4,0962026.04 | 81.7 | 60.5 | 73.3 | |
| Qwen2.5-VL-7BRelative inference throughput (Tp)=1.00×, Max visual tokens=4,0962026.04 | 80 | 75 | 78 | |
| LLaVA-1.5-7B + SD-RPNRelative inference throughput (Tp)=0.57×, Max visual tokens=4,0962026.04 | 70.4 | 71.1 | 70.7 | |
| LLaVA-1.5-7B + Q-ZoomRelative inference throughput (Tp)=0.58×, Max visual tokens=4,0962026.04 | 70.4 | 71.1 | 70.7 | |
| LLaVA-1.5-13B + Q-ZoomRelative inference throughput (Tp)=0.58×, Max visual tokens=4,0962026.04 | 61 | 64.5 | 61.8 | |
| LLaVA-1.5-13B + SD-RPNRelative inference throughput (Tp)=0.57×, Max visual tokens=4,0962026.04 | 60.9 | 65.8 | 62.8 | |
| LLaVA-1.5-7B + ViCropRelative inference throughput (Tp)=0.15×, Max visual tokens=4,0962026.04 | 53.9 | 50 | 52.4 | |
| LLaVA-1.5-7B + S2Relative inference throughput (Tp)=0.63×, Max visual tokens=4,0962026.04 | 53 | 59.1 | 55.5 | |
| LLaVA-1.5-7BRelative inference throughput (Tp)=1.00×, Max visual tokens=4,0962026.04 | 48.7 | 52.6 | 50.3 | |
| LLaVA-1.5-13B + ViCropRelative inference throughput (Tp)=0.21×, Max visual tokens=4,0962026.04 | 47.8 | 57.9 | 51.8 | |
| LLaVA-1.5-13BRelative inference throughput (Tp)=1.00×, Max visual tokens=4,0962026.04 | 47 | 56.6 | 50.8 | |
| LLaVA-1.5-13B + S2Relative inference throughput (Tp)=0.73×, Max visual tokens=4,0962026.04 | 43.5 | 59.2 | 49.7 |