Short-answer Visual Question Answering on RealWorldQA
65.1AccuracyQwen2.5-VL-3B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-VL-3BModel Category=Autoregressive Vision-Language Models2026.04 | 65.1 | |
| Fast-dVLMModel Category=Diffusion Vision-Language Models, Decoding Strategy=masked diffusion model decoding (MDM)2026.04 | 65.1 | |
| Fast-dVLMModel Category=Diffusion Vision-Language Models, Decoding Strategy=speculative decoding (spec.)2026.04 | 65.1 | |
| Intern-VL-2.5-4BModel Category=Autoregressive Vision-Language Models2026.04 | 64.6 | |
| LLaDA-VModel Category=Diffusion Vision-Language Models2026.04 | 63.2 | |
| MiniCPM-V-2 (3B)Model Category=Autoregressive Vision-Language Models2026.04 | 56.3 | |
| DimpleModel Category=Diffusion Vision-Language Models2026.04 | 55.4 | |
| LaViDaModel Category=Diffusion Vision-Language Models2026.04 | 54.5 | |
| VILA-1.5-3BModel Category=Autoregressive Vision-Language Models2026.04 | 53.2 |