Multi-discipline Multimodal Understanding on MMMU (val)
81.7Accuracygemini-2.5-pro-exp-03-25
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| gemini-2.5-pro-exp-03-252026.01 | 81.7 | — | — | |
| AgoraRouting=Pool of 5 baseline VLMs2026.01 | 79.2 | — | — | |
| Gemini-2.5-ProModel Category=Proprietary2026.02 | 78.78 | — | — | |
| OpenAI-o1Learning Protocol=Closed-source VLM2026.02 | 78.2 | — | — | |
| Claude-4-SonnetModel Category=Proprietary2026.02 | 74.44 | — | — | |
| Octopus-8B (Ours)Rollout count (rollout.n)=8, Generation Time (Gen.)=344.7, Total training time per step=958.12026.02 | 72.4 | — | — | |
| InternVL3-78BScale=78B2026.01 | 72.2 | — | — | |
| Guided-GRPOBackbone=Qwen3-8B, Training=G-GRPO+SFT2026.02 | 72.11 | — | — | |
| Qwen3-VL-8B-ThinkingLearning Protocol=Open-source Reasoning VLM2026.02 | 71.8 | — | — | |
| Qwen3-VL-8B-Instruct + DAPORollout count (rollout.n)=16, Learning Protocol=RLVR2026.02 | 71.4 | — | — | |
| InternVL3.5-8B-RLLearning Protocol=Open-source Reasoning VLM2026.02 | 71.2 | — | — | |
| Claude-3.7-SonnetLearning Protocol=Closed-source VLM2026.02 | 71 | — | — | |
| Qwen3-VL-8B-Instruct + GSPORollout count (rollout.n)=16, Generation Time (Gen.)=688.9, Total training time per step=1322.8, Learning Protocol=RLVR2026.02 | 70.8 | — | — | |
| gemini-2.0-flash2026.01 | 70.7 | — | — | |
| gpt-4o-2024-08-062026.01 | 70.7 | — | — | |
| Qwen3-VL-8B-Instruct + GRPORollout count (rollout.n)=16, Generation Time (Gen.)=679.1, Total training time per step=1428.4, Learning Protocol=RLVR2026.02 | 70.6 | — | — | |
| Qwen3-VL-32B-InstructContext Window=32K2026.03 | 70.6 | — | — | |
| Qwen3-VL-8B-Instruct + GRPORollout count (rollout.n)=8, Generation Time (Gen.)=364.9, Total training time per step=845.1, Learning Protocol=RLVR2026.02 | 70.4 | — | — | |
| qwen2.5vl-72b-instructScale=72B2026.01 | 70.2 | — | — | |
| Qwen3-VL-8B-Instruct + SRPORollout count (rollout.n)=16, Generation Time (Gen.)=816.6, Total training time per step=1543.3, Learning Protocol=RLVR2026.02 | 69.7 | — | — | |
| Qwen3-VL-8B-Instruct + SRPORollout count (rollout.n)=8, Generation Time (Gen.)=410.9, Total training time per step=895.8, Learning Protocol=RLVR2026.02 | 69.4 | — | — | |
| GPT-4oZero-shot=true, Decoding Strategy=greedy2024.09 | 69.2 | — | — | |
| GPT-4o2024.08 | 69.1 | — | — | |
| GPT-4o2024.12 | 69.1 | — | — | |
| GPT-4oLearning Protocol=Closed-source VLM2026.02 | 69.1 | — | — | |
| Qwen3-VL-32B-InstructContext Window=4K2026.03 | 68.6 | — | — | |
| Qwen3-VL-8B-Instruct + GSPORollout count (rollout.n)=8, Generation Time (Gen.)=361.7, Total training time per step=753.1, Learning Protocol=RLVR2026.02 | 67.4 | — | — | |
| GPT-4oModel Category=Proprietary2026.02 | 67.33 | — | — | |
| MiMo-VL-7B-RLLearning Protocol=Open-source Reasoning VLM2026.02 | 66.7 | — | — | |
| gemma-3-27bScale=27B2026.01 | 64.9 | — | — | |
| MiMo-VL-7B-SFTLearning Protocol=Open-source Reasoning VLM2026.02 | 64.6 | — | — | |
| Qwen3-VL-8B-InstructContext Window=32K2026.03 | 64.6 | — | — | |
| Base (Qwen3-8B)Backbone=Qwen3-8B2026.02 | 62.44 | — | — | |
| Gemini-1.5-Pro2024.12 | 62.2 | — | — | |
| Qwen3-VL-8B-InstructLearning Protocol=Base VLM2026.02 | 62 | — | — | |
| Vision-R1-32BModel Category=Open-Source2026.02 | 61.33 | — | — | |
| Qwen2.5-VL-32BModel Category=Open-Source2026.02 | 61 | — | — | |
| Qwen3-VL-8B-InstructContext Window=4K2026.03 | 60.7 | — | — | |
| Gemini-1.5-ProZero-shot=true, Decoding Strategy=greedy2024.09 | 60.6 | — | — | |
| gpt-4o-mini2026.01 | 60 | — | — | |
| Gemini Ultra 1.0open-source=false2024.04 | 59.4 | — | — | |
| Claude-3 Opusopen-source=false2024.04 | 59.4 | — | — | |
| InternVL2.5-26BModel Category=Open-Source2026.02 | 59.33 | — | — | |
| qwen2.5vl-7b-instructScale=7B2026.01 | 58.6 | — | — | |
| Lumina-DiMOOModel Architecture Style=Diff. Based2026.03 | 58.6 | — | — | |
| Gemini Pro 1.5open-source=false2024.04 | 58.5 | — | — | |
| MMR1-32BModel Category=Open-Source2026.02 | 57.89 | — | — | |
| GPT-4VLLM=Private, Zero-shot=true2024.03 | 56.8 | — | — | |
| GPT-4Vopen-source=false2024.04 | 56.8 | — | — | |
| LLaVA-OneVision-72BModel Scale=72B2024.08 | 56.8 | — | — | |
| GPT-4VVersion=V-Preview2024.08 | 56.8 | — | — | |
| GPT-4V2024.12 | 56.8 | — | — | |
| VL-Rethinker-7BModel Category=Open-Source2026.02 | 56.67 | — | — | |
| Qwen3-VL-4BModel=Qwen3-VL, Parameter Count=4B2026.04 | 55.78 | — | — | |
| BAGELModel Architecture Style=AR Based2026.03 | 55.3 | — | — | |
| Phi-4-reasoning-vision-15B2026.03 | 54.3 | — | — | |
| Qwen2-VL-InstructLLM=Qwen2-7B, Vision Encoder=DFN-CLIP-H2024.12 | 54.1 | — | — | |
| Qwen2-VLSize=7B, #token /image tile=Native resolution (unfixed)2024.12 | 54.1 | — | — | |
| GPT-4VZero-shot=true, Decoding Strategy=greedy2024.09 | 53.8 | — | — | |
| Claude-3 Sonnetopen-source=false2024.04 | 53.1 | — | — | |
| LLaVA-v1.5-13BModel Category=Open-Source2026.02 | 53 | — | — | |
| Phi-4-reasoning-vision-15BProtocol=force nothink2026.03 | 52 | — | — | |
| Kimi-VL-A3B-Instruct2026.03 | 52 | — | — | |
| InternVL 1.2#param=40B, open-source=false2024.04 | 51.6 | — | — | |
| Qwen-VL-MaxModel Category=Proprietary2026.02 | 51.44 | — | — | |
| Qwen-VL-Maxopen-source=false2024.04 | 51.3 | — | — | |
| Phi-3-vision-128kModel Category=Open-Source2026.02 | 51.11 | — | — | |
| Qwen2.5-VL-3BModel=Qwen2.5-VL, Parameter Count=3B2026.04 | 51.11 | — | — | |
| LLaVA-NeXTLLM=Hermes-2-Yi-34B, Resolution=672, Zero-shot=true2024.03 | 51.1 | — | — | |
| LLaVA-NeXT#param=35B, open-source=false2024.04 | 51.1 | — | — | |
| LLaVA-NeXT-34BModel Scale=30B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 51.1 | — | — | |
| PVC InternVL2 (Ours)Size=8B, #token /image tile=642024.12 | 50.9 | — | — | |
| InternVL 3.5 2BModel Size=2B2026.03 | 50.51 | — | — | |
| SpB2.0-VL-5BModel=SpB2.0-VL, Parameter Count=5B2026.04 | 50.33 | — | — | |
| Claude-3 Haikuopen-source=false2024.04 | 50.2 | — | — | |
| gemma-3-12b-it2026.03 | 50 | — | — | |
| Step-1V#param=100B, open-source=false2024.04 | 49.9 | — | — | |
| Cambrian-34BModel Scale=30B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 49.7 | — | — | |
| InternVL2Size=8B, #token /image tile=2562024.12 | 49.3 | — | — | |
| LLaVA-OneVision-7BModel=LLaVA-OneVision, Parameter Count=7B2026.04 | 49 | — | — | |
| Show-o2Model Architecture Style=AR Based2026.03 | 48.9 | — | — | |
| LLaVA-OneVision-7BModel Scale=7B2024.08 | 48.8 | — | — | |
| LLaVA-OVLLM=Qwen2-7B, Vision Encoder=SigLIP-SO400M2024.12 | 48.8 | — | — | |
| LLaVA-OVSize=7B, #token /image tile=7292024.12 | 48.8 | — | — | |
| Mini-GeminiLLM=Hermes-2-Yi-34B, Resolution=336, Zero-shot=true2024.03 | 48.7 | — | — | |
| LLaDA-VModel Architecture Style=Diff. Based2026.03 | 48.6 | — | — | |
| Xuanwu VL-2BModel Size=2B2026.03 | 48.11 | — | — | |
| Mini-Gemini-HDLLM=Hermes-2-Yi-34B, Resolution=672, Zero-shot=true2024.03 | 48 | — | — | |
| Mini-Gemini#param=35B, open-source=false2024.04 | 48 | — | — | |
| InternVL2-4BModel=InternVL2, Parameter Count=4B, Result Source=OpenVLM Leaderboard (OpenCompass, 2025)2026.04 | 48 | — | — | |
| Gemini ProLLM=Private, Zero-shot=true2024.03 | 47.9 | — | — | |
| Gemini Pro 1.0open-source=false2024.04 | 47.9 | — | — | |
| MM1.5-30BModel Scale=30B, Zero-shot=true, Decoding Strategy=greedy2024.09 | 47.4 | — | — | |
| LLaVA-OV (SI)LLM=Qwen2-7B, Vision Encoder=SigLIP-SO400M2024.12 | 47.3 | — | — | |
| Llama-3.2-11BModel Category=Open-Source2026.02 | 46.89 | — | — | |
| InternVL 3.0 2BModel Size=2B2026.03 | 45.89 | — | — | |
| Qwen-VL-PlusLLM=Private, Zero-shot=true2024.03 | 45.2 | — | — | |
| Qwen-VL-Plusopen-source=false2024.04 | 45.2 | — | — | |
| InternVL 1.5#param=26B, open-source=false2024.04 | 45.2 | — | — | |
| DimpleModel Architecture Style=Diff. Based2026.03 | 45.2 | — | — |