Real-world Multimodal Understanding on MME-RealWorld Lite
54.9Lite ScoreZwZ-Qwen3-VL-4B + Q-Zoom
Evaluation Results
| Method | Links | |
|---|---|---|
| ZwZ-Qwen3-VL-4B + Q-ZoomRelative inference throughput (Tp)=0.66×2026.04 | 54.9 | |
| ZwZ-Qwen3-VL-4BRelative inference throughput (Tp)=1.00×2026.04 | 54.3 | |
| Qwen2.5-VL-7B + ThymeRelative inference throughput (Tp)=0.21×2026.04 | 53.7 | |
| Qwen3-VL-4B + Q-ZoomRelative inference throughput (Tp)=0.73×2026.04 | 53.5 | |
| ZwZ-Qwen2.5-VL-7B + Q-ZoomRelative inference throughput (Tp)=0.76×2026.04 | 53.2 | |
| ZwZ-Qwen2.5-VL-7BRelative inference throughput (Tp)=1.00×2026.04 | 52.5 | |
| Qwen3-VL-4B + SD-RPNRelative inference throughput (Tp)=0.58×2026.04 | 51.8 | |
| Qwen2.5-VL-7B + DeepEyesEvaluation protocol=Directly cited (†)2026.04 | 50.9 | |
| Qwen2.5-VL-7B + Q-ZoomRelative inference throughput (Tp)=0.86×2026.04 | 48 | |
| Qwen3-VL-4BRelative inference throughput (Tp)=1.00×2026.04 | 47.4 | |
| Qwen2.5-VL-7B + SD-RPNRelative inference throughput (Tp)=0.77×2026.04 | 46.4 | |
| Qwen2.5-VL-3B + Q-ZoomRelative inference throughput (Tp)=0.67×2026.04 | 44 | |
| Qwen2.5-VL-7BRelative inference throughput (Tp)=1.00×2026.04 | 42.7 | |
| Qwen2.5-VL-3B + SD-RPNRelative inference throughput (Tp)=0.66×2026.04 | 41.9 | |
| Qwen2.5-VL-3BRelative inference throughput (Tp)=1.00×2026.04 | 41.6 | |
| LLaVA-1.5-13B + S2Relative inference throughput (Tp)=0.73×2026.04 | 36.3 | |
| LLaVA-1.5-13B + SD-RPNRelative inference throughput (Tp)=0.57×2026.04 | 31.4 | |
| LLaVA-1.5-13B + Q-ZoomRelative inference throughput (Tp)=0.58×2026.04 | 31.3 | |
| LLaVA-1.5-13B + ViCropRelative inference throughput (Tp)=0.21×2026.04 | 30.1 | |
| LLaVA-1.5-7B + S2Relative inference throughput (Tp)=0.63×2026.04 | 28.5 | |
| LLaVA-1.5-13BRelative inference throughput (Tp)=1.00×2026.04 | 27.8 | |
| LLaVA-1.5-7BRelative inference throughput (Tp)=1.00×2026.04 | 27.7 | |
| LLaVA-1.5-7B + SD-RPNRelative inference throughput (Tp)=0.57×2026.04 | 27.7 | |
| LLaVA-1.5-7B + Q-ZoomRelative inference throughput (Tp)=0.58×2026.04 | 27.7 | |
| LLaVA-1.5-7B + ViCropRelative inference throughput (Tp)=0.15×2026.04 | 27.6 |