Multimodal Reasoning on MMMU Pro (Accuracy)
85.6AccuracyCoT2-Meta
Evaluation Results
| Method | Links | |
|---|---|---|
| CoT2-MetaStrategy=Ours (CoT2-Meta), Inference Budget=C=16, Textification=OCR+Captioning2026.03 | 85.6 | |
| ReST-MCTS*Strategy=ReST-MCTS*, Inference Budget=C=16, Textification=OCR+Captioning2026.03 | 81.3 | |
| Vanilla ToTStrategy=Vanilla ToT, Inference Budget=C=16, Textification=OCR+Captioning2026.03 | 77.8 | |
| Gemini-2.5 (Pro)Model Tier=Pro2026.01 | 76.96 | |
| Qwen3.5Parameters=35BA3B, Internal Reasoning (Think mode)=false2026.05 | 75.1 | |
| Gemma4Parameters=26BA4B, Internal Reasoning (Think mode)=false2026.05 | 73.8 | |
| Best-of-16Strategy=Best-of-16, Inference Budget=C=16, Textification=OCR+Captioning2026.03 | 73.1 | |
| SenseNova-U1Parameters=30BA3B, Internal Reasoning (Think mode)=true2026.05 | 72.83 | |
| Qwen3-VL (Thinking)Thinking Mode=true, Parameters=235B-A22B2026.01 | 72.37 | |
| Seed-1.5-VL (Thinking)Thinking Mode=true2026.01 | 70.6 | |
| Qwen3.5Parameters=9B, Internal Reasoning (Think mode)=false2026.05 | 70.1 | |
| GPT-4oActivation Replay=false2025.11 | 69.1 | |
| Greedy CoTStrategy=Greedy CoT, Inference Budget=C=16, Textification=OCR+Captioning2026.03 | 68.4 | |
| Claude-3.5 SonnetActivation Replay=false2025.11 | 68.3 | |
| SenseNova-U1Parameters=8B, Internal Reasoning (Think mode)=true2026.05 | 67.69 | |
| STEP3-VL-10B (PaCoRe)Reasoning Strategy=PaCoRe, Parameters=10B2026.01 | 67.18 | |
| GLM-4.6VParameters=106B-A12B2026.01 | 65.84 | |
| STEP3-VL-10B (SeRe)Reasoning Strategy=SeRe, Parameters=10B2026.01 | 64.08 | |
| Qwen3VLParameters=30BA3B, Internal Reasoning (Think mode)=true2026.05 | 63 | |
| OpenAI-o1#Data=/2025.06 | 62.4 | |
| Qwen3VLParameters=8B, Internal Reasoning (Think mode)=true2026.05 | 60.4 | |
| LongCat-NextParameters=68BA3B, Internal Reasoning (Think mode)=false2026.05 | 60.3 | |
| GPT-4oModel Category=Closed-Source MLLMs2026.01 | 56.13 | |
| DLRModel Type=Open-Source, Backbone=Qwen3-VL-8B-Thinking2026.04 | 56.1 | |
| LVRModel Type=Open-Source, Backbone=Qwen3-VL-8B-Thinking2026.04 | 55.3 | |
| Qwen3-VLModel Size=4B, Training Stage=instruct2026.05 | 53.2 | |
| PixelReasonerModel Type=Open-Source, Backbone=Qwen3-VL-8B-Thinking2026.04 | 53.1 | |
| GPT-4oSize=—2026.01 | 51.9 | |
| GPT-4o#Data=/2025.06 | 51.9 | |
| GPT-4oModel Type=Proprietary Model2026.04 | 51.9 | |
| Gemini-2.0-FlashActivation Replay=false2025.11 | 51.7 | |
| Claude-3.7-Sonnet#Data=/2025.06 | 51.5 | |
| Qwen3-VLModel Size=4B, Training Stage=S-12026.05 | 51.5 | |
| Gemini-1.5-ProModel Category=Closed-Source MLLMs2026.01 | 51.47 | |
| Qwen3-VL-SegModel Size=4B, Training Stage=S-22026.05 | 51.3 | |
| MM-Eureka-Qwen-32BActivation Replay=true2025.11 | 51 | |
| Qwen3-VL-8B-ThinkingModel Type=Open-Source, Backbone=Qwen3-VL-8B-Thinking2026.04 | 50.2 | |
| ICoTModel Type=Open-Source, Backbone=Qwen3-VL-8B-Thinking2026.04 | 49.6 | |
| MM-Eureka-Qwen-32BActivation Replay=false2025.11 | 49.6 | |
| Qwen2.5-VL-72B-IT#Data=/2025.06 | 49.5 | |
| Qwen3-VL-8B-Instruct + CAREBackbone=Qwen3-VL-8B-Instruct, Method=CARE2025.12 | 46.7 | |
| Kimi-VL-16BActivation Replay=false2025.11 | 46.3 | |
| MiMo-VL-7B-RLModel Type=Reasoning2025.12 | 46.2 | |
| MiMo-VL-7B-SFTModel Type=Instruct2025.12 | 45.2 | |
| VL-Rethinker-7BActivation Replay=true2025.11 | 42.7 | |
| Perception-R1-7B#Data=1.4K2025.06 | 42.4 | |
| PRCO-7BBackbone=Qwen2.5-VL-7B2026.03 | 42.08 | |
| Qwen3-VL-8B-Instruct + CAREBackbone=Qwen3-VL-8B-Instruct, Method=CARE2025.12 | 41.7 | |
| VL-Rethinker-7BActivation Replay=false2025.11 | 41.7 | |
| PAPO-D-7BBackbone=Qwen2.5-VL-7B2026.03 | 41.5 | |
| DAPOBackbone=Qwen2.5-VL-7B2026.03 | 41.38 | |
| Vision-SR1-7BBackbone=Qwen2.5-VL-7B2026.03 | 41.38 | |
| MMR1-Math-v0-7BActivation Replay=false2025.11 | 41.3 | |
| MMR1-Math-v0-7BActivation Replay=true2025.11 | 41.3 | |
| Ours [with DPS and annealing]Model Category=Reasoning MLLMs, Training Strategy=Two-stage RL2026.01 | 41 | |
| InternVL3.5Parameters=8B2026.04 | 41 | |
| AdaSteerModel=QWEN3.52026.06 | 41 | |
| Ours [with DPS]Model Category=Reasoning MLLMs, Training Strategy=DPS2026.01 | 40.7 | |
| OursModel Category=Reasoning MLLMs, Training Strategy=DAPO2026.01 | 40.6 | |
| MiMo-VL-7B-RLModel Type=Reasoning2025.12 | 40.3 | |
| PAPO-G-7BBackbone=Qwen2.5-VL-7B2026.03 | 40.11 | |
| MM-Eureka-Qwen-7BActivation Replay=true2025.11 | 40 | |
| Qwen2.5-VL-7B + CAREBackbone=Qwen2.5-VL-7B, Method=CARE2025.12 | 39.7 | |
| VPPO-7BBackbone=Qwen2.5-VL-7B2026.03 | 39.65 | |
| AdaSteerModel=QWEN3-VL2026.06 | 39.6 | |
| VLAA-Thinker 7BModel Category=Reasoning MLLMs2026.01 | 39.5 | |
| MM-Eureka-Qwen-7BActivation Replay=false2025.11 | 39.5 | |
| MiMo-VL-7B-SFTModel Type=Instruct2025.12 | 39.4 | |
| Zero-shotModel=QWEN3-VL2026.06 | 39.3 | |
| ECSOModel=QWEN3-VL2026.06 | 39.3 | |
| MARSModel=QWEN3-VL2026.06 | 39.3 | |
| GRPOBackbone=Qwen2.5-VL-7B2026.03 | 39.01 | |
| Qwen2.5-VL-7B + GSPOBackbone=Qwen2.5-VL-7B, RL Method=GSPO2025.12 | 38.9 | |
| SophiaVL-R1-7B#Data=130K2025.06 | 38.8 | |
| R1-ShareVL-7BBackbone=Qwen2.5-VL-7B2026.03 | 38.32 | |
| MM-Eureka-7B#Data=15K2025.06 | 38.3 | |
| Qwen2.5-VL-7BActivation Replay=false2025.11 | 38.3 | |
| Perception-R1-7BBackbone=Qwen2.5-VL-7B2026.03 | 38.2 | |
| InternVL3.5Parameters=4B2026.04 | 38.2 | |
| Qwen2.5VL 7BModel Category=Open-Source General MLLMs2026.01 | 38 | |
| MixedR1 7BModel Category=Reasoning MLLMs2026.01 | 38 | |
| OpenVLThinker-7B#Data=25K2025.06 | 37.8 | |
| SASAModel=QWEN3-VL2026.06 | 37.8 | |
| Vision-R1-7B#Data=200K2025.06 | 37.6 | |
| BARD-VLParameters=8B, B=42026.04 | 37.6 | |
| LLaVA-OneVision-1.5 8BModel Type=Instruct2025.12 | 37.4 | |
| Qwen2.5-VL-7B + DAPOBackbone=Qwen2.5-VL-7B, RL Method=DAPO2025.12 | 37.3 | |
| VLAA-Thinker-7B#Data=25K2025.06 | 37.2 | |
| Qwen2.5-VL-7B + CAREBackbone=Qwen2.5-VL-7B, Method=CARE2025.12 | 37.1 | |
| Vision-Matters-7BBackbone=Qwen2.5-VL-7B2026.03 | 37.1 | |
| Qwen2.5-VL-7B-IT#Data=/2025.06 | 37 | |
| Qwen3-VLParameters=2B2026.03 | 36.5 | |
| Zero-shotModel=INTERNVL 3.52026.06 | 36.5 | |
| ECSOModel=INTERNVL 3.52026.06 | 36.5 | |
| MARSModel=INTERNVL 3.52026.06 | 36.5 | |
| Qwen2.5-VL-7B + GRPOBackbone=Qwen2.5-VL-7B, RL Method=GRPO2025.12 | 36.4 | |
| Qwen2.5-VL-7B + GSPOBackbone=Qwen2.5-VL-7B, RL Method=GSPO2025.12 | 36.4 | |
| Qwen2.5-VL-7BModel Type=Instruct2025.12 | 36.3 | |
| NoisyRollout-7BBackbone=Qwen2.5-VL-7B2026.03 | 36.3 | |
| AdaSteerModel=INTERNVL 3.52026.06 | 36.1 |