Multi-modal Reasoning on MMVet (test)
80.8AccuracyGPT-4o
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| GPT-4oModel Category=Closed-Source Models2026.01 | 80.8 | — | |
| MMCTAgentsetup=zero-shot, critic=true2024.05 | 74.24 | — | |
| Qwen-VL-MaxModel Category=Closed-Source Models2026.01 | 73.2 | — | |
| R1-OneVisionBase Model Scale=7B, Base Architecture=Qwen2.5-VL2026.01 | 71.6 | 440.7 | |
| FAST-7BBase Model Scale=7B, Base Architecture=Qwen2.5-VL2026.01 | 71.2 | 114.1 | |
| GPRO-7BBase Model Scale=7B, Base Architecture=Qwen2.5-VL2026.01 | 70.9 | 118.8 | |
| MM-R1Base Model Scale=7B, Base Architecture=Qwen2.5-VL2026.01 | 70.6 | 137.9 | |
| MMCTAgentsetup=zero-shot, critic=false2024.05 | 70.51 | — | |
| Claude-3.5 SonnetModel Category=Closed-Source Models2026.01 | 68.7 | — | |
| OpenVLThinkerBase Model Scale=7B, Base Architecture=Qwen2.5-VL2026.01 | 68.5 | 312.7 | |
| Qwen2.5-VL-7BBase Model Scale=7B, Base Architecture=Qwen2.5-VL2026.01 | 67.1 | 132.5 | |
| LMM-R1Base Model Scale=3B, Base Architecture=Qwen2.5-VL2026.01 | 65.9 | 166.3 | |
| GPRO-3BBase Model Scale=3B, Base Architecture=Qwen2.5-VL2026.01 | 65.2 | 108.4 | |
| Gemini 1.5 Prosetup=zero-shot2024.05 | 64.2 | — | |
| FAST-3BBase Model Scale=3B, Base Architecture=Qwen2.5-VL2026.01 | 64 | 112.7 | |
| Qwen2-VL-7BBase Model Scale=7B, Base Architecture=Qwen2-VL2026.01 | 62 | 132.5 | |
| Curr-ReFTBase Model Scale=3B, Base Architecture=Qwen2.5-VL2026.01 | 62 | 117.6 | |
| VideoAuto-R1Backbone=Qwen2.5-VL-7B2026.01 | 61.9 | — | |
| Qwen2.5-VL-3BBase Model Scale=3B, Base Architecture=Qwen2.5-VL2026.01 | 61.3 | 138.8 | |
| GPT-4Vsetup=zero-shot2024.05 | 60.2 | — | |
| Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B2026.01 | 60 | — | |
| CNTPModel=LLaVA-CoT2025.07 | 58.5 | — | |
| Stochastic DecodingModel=Llama-3.2-11B-Vision-Instruct2025.07 | 53.5 | — | |
| CNTPModel=Llama-3.2-11B-Vision-Instruct2025.07 | 53.5 | — | |
| Stochastic DecodingModel=LLaVA-CoT2025.07 | 53 | — | |
| Claude 3 Opussetup=zero-shot2024.05 | 51.7 | — | |
| Claude 3 Sonnetsetup=zero-shot2024.05 | 51.3 | — | |
| Greedy DecodingModel=Llama-3.2-11B-Vision-Instruct2025.07 | 48 | — | |
| Greedy DecodingModel=LLaVA-CoT2025.07 | 47.7 | — | |
| MulberryBase Model Scale=7B, Base Architecture=Qwen2-VL2026.01 | 43.9 | 218.3 | |
| FullTraining Data Fraction=100%2026.05 | 30.9 | — | |
| MAGICTraining Data Fraction=20%2026.05 | 29.8 | — | |
| ICONSTraining Data Fraction=20%2026.05 | 29.7 | — | |
| RandomTraining Data Fraction=20%2026.05 | 29.5 | — | |
| CC12M Split 3Partitioning=Clustering2026.04 | 28.4 | — | |
| PivotMergePartitioning=Clustering2026.04 | 27.8 | — | |
| TIES MergingPartitioning=Clustering2026.04 | 27.3 | — | |
| Weight AveragePartitioning=Clustering2026.04 | 26.9 | — | |
| MetaGPTPartitioning=Clustering2026.04 | 26.6 | — | |
| Self-FilterTraining Data Fraction=20%2026.05 | 26.6 | — | |
| CC12M Split 1Partitioning=Clustering2026.04 | 26.3 | — | |
| TIES w/ DAREPartitioning=Clustering2026.04 | 24.3 | — | |
| TSV-MPartitioning=Clustering2026.04 | 23.2 | — | |
| Task ArithmeticPartitioning=Clustering2026.04 | 21.6 | — | |
| EL2NTraining Data Fraction=20%2026.05 | 21.1 | — | |
| CC12M Split 5Partitioning=Clustering2026.04 | 20.6 | — | |
| CC12M Split 2Partitioning=Clustering2026.04 | 20 | — | |
| CC12M Split 4Partitioning=Clustering2026.04 | 17.5 | — | |
| Iso-CPartitioning=Clustering2026.04 | 8.2 | — |