Visual Reasoning on MM-Vet
82.7ScoreUniDFlow
Evaluation Results
| Method | Links | |
|---|---|---|
| UniDFlowParams=4B2026.02 | 82.7 | |
| MammothModa-2Params=4B2026.02 | 79.4 | |
| MudditParams=4B2026.02 | 76.2 | |
| EMMAParams=4B2026.02 | 73 | |
| Qwen3-VLParams=4B2026.02 | 72.5 | |
| GPT-4oZero-shot=true, Decoding Strategy=greedy2024.09 | 69.1 | |
| BAGELParams=7B2026.02 | 67.2 | |
| Gemini-1.5-ProZero-shot=true, Decoding Strategy=greedy2024.09 | 64 | |
| DeepSeek-VL2Params=4B2026.02 | 62.8 | |
| Qwen2.5-VLParams=3B2026.02 | 61.8 | |
| LLaVA-NeXT-34BModel Scale=30B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 57.4 | |
| GPT-4VZero-shot=true, Decoding Strategy=greedy2024.09 | 56.8 | |
| MM1.5-30BModel Scale=30B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 52 | |
| Janus-ProParams=7B2026.02 | 50 | |
| MM1-30BModel Scale=30B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 48.7 | |
| Pretrained-gpt2-encoder-LLaVA1.6-13BLLM Backbone=Vicuna-13B, Vision Encoder Initialization=Pretrained (Language Bias), Model Scale=13B2026.04 | 48.7 | |
| Phi-3-Vision-4BModel Scale=3B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 46.2 | |
| Scratch-gpt2-encoder-LLaVA1.6-13BLLM Backbone=Vicuna-13B, Vision Encoder Initialization=Scratch, Model Scale=13B2026.04 | 44.3 | |
| LLaVA-NeXT-7BModel Scale=7B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 43.9 | |
| MM1-3BModel Scale=3B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 43.7 | |
| MM1.5-3B-MoEModel Scale=3B, Architecture=MoE, Zero-shot=true, Decoding Strategy=greedy2024.09 | 43.7 | |
| Pretrained-gpt2-encoder-LLaVA1.6-7BLLM Backbone=Mistral-7B, Vision Encoder Initialization=Pretrained (Language Bias), Model Scale=7B2026.04 | 43.2 | |
| MM1.5-7BModel Scale=7B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 42.2 | |
| MM1-7BModel Scale=7B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 42.1 | |
| MM1.5-3BModel Scale=3B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 41 | |
| Scratch-gpt2-encoder-LLaVA1.6-7BLLM Backbone=Mistral-7B, Vision Encoder Initialization=Scratch, Model Scale=7B2026.04 | 40.9 | |
| TokenFlow-XLParams=13B2026.02 | 40.7 | |
| MM1.5-1B-MoEModel Scale=1B, Architecture=MoE, Zero-shot=true, Decoding Strategy=greedy2024.09 | 39.8 | |
| InternVL2-2BModel Scale=1B, Zero-shot=true, Decoding Strategy=beam search2024.09 | 39.7 | |
| MM1-1BModel Scale=1B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 39.4 | |
| MiniCPM-V 2.0-3BModel Scale=3B, Zero-shot=true, Decoding Strategy=beam search2024.09 | 38.2 | |
| MM1.5-1BModel Scale=1B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 37.4 | |
| DeepSeek-VLModel Scale=1B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 34.8 | |
| TinyLLaVAModel Scale=3B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 32 | |
| VILA-UParams=7B2026.02 | 27.7 | |
| TinyLLaVAModel Scale=1B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 25.8 | |
| SPHINX-TinyModel Scale=1B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 23.8 | |
| Pretrained-gpt2-encoder-LLaVA1.5-7BLLM Backbone=Vicuna-7B, Vision Encoder Initialization=Pretrained (Language Bias), Model Scale=7B2026.04 | 18.5 | |
| Scratch-gpt2-encoder-LLaVA1.5-7BLLM Backbone=Vicuna-7B, Vision Encoder Initialization=Scratch, Model Scale=7B2026.04 | 11.1 | |
| ChameleonParams=7B2026.02 | 8.3 |