Long-answer Visual Question Answering on MMMU-Pro-V
26.3AccuracyQwen2.5-VL-3B
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-VL-3BModel Category=Autoregressive Vision-Language Models, Tokens/NFE=1.002026.04 | 26.3 | |
| Intern-VL-2.5-4BModel Category=Autoregressive Vision-Language Models, Tokens/NFE=1.002026.04 | 24.6 | |
| Fast-dVLMModel Category=Diffusion Vision-Language Models, Decoding Strategy=speculative decoding (spec.), Tokens/NFE=2.632026.04 | 24.6 | |
| Fast-dVLMModel Category=Diffusion Vision-Language Models, Decoding Strategy=masked diffusion model decoding (MDM), Tokens/NFE=1.952026.04 | 21.4 | |
| LLaDA-VModel Category=Diffusion Vision-Language Models, Tokens/NFE=1.002026.04 | 18.6 | |
| DimpleModel Category=Diffusion Vision-Language Models, Tokens/NFE=1.002026.04 | 12.4 | |
| LaViDaModel Category=Diffusion Vision-Language Models, Tokens/NFE=1.002026.04 | 10.5 | |
| MiniCPM-V-2 (3B)Model Category=Autoregressive Vision-Language Models, Tokens/NFE=1.002026.04 | 10.3 | |
| VILA-1.5-3BModel Category=Autoregressive Vision-Language Models, Tokens/NFE=1.002026.04 | 6.1 |