Multi-modal Reasoning on EMMA
38.5AccuracyQwen2.5-VL-Instruct
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen2.5-VL-InstructSize=72B2026.05 | 38.5 | — | — | |
| Claude-3.5Size=-2026.05 | 38.1 | — | — | |
| AnE-3rdTraining Stage=Round 32026.05 | 34.6 | — | — | |
| AnE-2ndTraining Stage=Round 22026.05 | 33 | — | — | |
| GPT-4oModel Category=Closed-Source MLLMs2026.01 | 32.7 | — | — | |
| GPT-4oSize=-2026.05 | 32.7 | — | — | |
| PTA-GRPOBase Model=Qwen2.5-7B-VL2025.10 | 31.9 | — | — | |
| STRIDEBase model=Qwen2.5-VL-7B2026.06 | 31.5 | — | — | |
| MoCASize=7B2026.05 | 31.3 | — | — | |
| ReLaX-VL-7BModel Category=Reasoning Multimodal LLM, Parameters=7B2025.12 | 30.6 | — | — | |
| AnE-1stTraining Stage=Round 12026.05 | 30.5 | — | — | |
| Preliminary RLTraining Phase=Warm-up2026.05 | 30.1 | — | — | |
| VL-Rethinker-7BModel Category=Reasoning Multimodal LLM, Parameters=7B2025.12 | 29.7 | — | — | |
| VL-RethinkerSize=7B2026.05 | 29.7 | — | — | |
| VL-RethinkerEvaluation Source=original paper2026.05 | 29.7 | — | — | |
| SRPO-7BModel Category=Reasoning Multimodal LLM, Parameters=7B2025.12 | 29.6 | — | — | |
| SRPOEvaluation Source=original paper2026.05 | 29.6 | — | — | |
| SRPOBase Model=Qwen2.5-7B-VL2025.10 | 29.6 | — | — | |
| SRPOBase model=Qwen2.5-VL-7B2026.06 | 29.6 | — | — | |
| DAPO (Ours)Model Category=Multi-image/Video Enhancing MLLMs, Training Strategy=DAPO2026.01 | 29.3 | — | — | |
| Two-stage RL (Ours)Model Category=Multi-image/Video Enhancing MLLMs, Training Strategy=Two-stage RL (DPS and annealing)2026.01 | 28.6 | — | — | |
| Metis-RISEEvaluation Source=reproduced2026.05 | 28.5 | — | — | |
| SPARK-VL-7BEvaluation Source=reproduced2026.05 | 28.5 | — | — | |
| DPS (Ours)Model Category=Multi-image/Video Enhancing MLLMs, Training Strategy=DPS2026.01 | 28.4 | — | — | |
| LLaVA-Critic-R1Evaluation Source=original paper2026.05 | 28.3 | — | — | |
| GPT-4o-miniSize=-2026.05 | 27.3 | — | — | |
| ReLaX-VL-3BModel Category=Reasoning Multimodal LLM, Parameters=3B2025.12 | 26.9 | — | — | |
| OpenVLThinkerEvaluation Source=original paper2026.05 | 26.8 | — | — | |
| OpenVLThinker-7BModel Category=Reasoning Multimodal LLM, Parameters=7B2025.12 | 26.6 | — | — | |
| mPLUG-Owl3Model Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=8B2026.01 | 24.8 | — | — | |
| Qwen2.5-VL-7B-InstructEvaluation Source=original paper2026.05 | 24.6 | — | — | |
| OpenMMReasoner-7BEvaluation Source=reproduced2026.05 | 24.5 | — | — | |
| Vision-R1Evaluation Source=reproduced2026.05 | 23.6 | — | — | |
| MM-Eureka-7BModel Category=Reasoning Multimodal LLM, Parameters=7B2025.12 | 23.5 | — | — | |
| MM-EurekaBase model=Qwen2.5-VL-7B2026.06 | 23.5 | — | — | |
| TW-GRPOModel Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=7B2026.01 | 22.5 | — | — | |
| Vision-R1-7BModel Category=Reasoning Multimodal LLM, Parameters=7B2025.12 | 22.4 | — | — | |
| Qwen2.5-VL-7BModel Category=General Multimodal LLM, Parameters=7B2025.12 | 21.5 | — | — | |
| MM-Eureka-8BModel Category=Reasoning Multimodal LLM, Parameters=8B2025.12 | 21.5 | — | — | |
| Qwen2.5-VL-InstructSize=7B2026.05 | 21.5 | — | — | |
| BaseBase Model=Qwen2.5-7B-VL2025.10 | 21.5 | — | — | |
| InternVL2.5Model Category=Open-Source General MLLMs, Parameter Scale=8B2026.01 | 21 | — | — | |
| Intern2.5-VL-8BModel Category=General Multimodal LLM, Parameters=8B2025.12 | 20.6 | — | — | |
| Qwen2.5-VLModel Category=Open-Source General MLLMs, Parameter Scale=7B2026.01 | 20.4 | — | — | |
| Mantis-Idefics2Model Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=8B2026.01 | 20.3 | — | — | |
| Qwen2-VL-7BModel Category=General Multimodal LLM, Parameters=7B2025.12 | 20.2 | — | — | |
| BaseBase model=Qwen2.5-VL-7B2026.06 | 20.2 | — | — | |
| Intern2-VL-8BModel Category=General Multimodal LLM, Parameters=8B2025.12 | 19.8 | — | — | |
| Pixel ReasonerSize=7B2026.05 | 19.8 | — | — | |
| Revisual-R1Evaluation Source=reproduced2026.05 | 19.5 | — | — | |
| Qwen2.5-VL-3BModel Category=General Multimodal LLM, Parameters=3B, reproduced=true2025.12 | 19.2 | — | — | |
| LLaVA-NeXT-InterleaveModel Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=7B2026.01 | 19 | — | — | |
| MMR1Evaluation Source=reproduced2026.05 | 18.6 | — | — | |
| Llava-OV-7BModel Category=General Multimodal LLM, Parameters=7B2025.12 | 18.3 | — | — | |
| Llava-OVSize=7B2026.05 | 18.3 | — | — | |
| DeepEyesSize=7B2026.05 | 18.1 | — | — | |
| VideoRFTModel Category=Multi-image/Video Enhancing MLLMs, Parameter Scale=7B2026.01 | 17.8 | — | — | |
| mPLUG-Owl3Size=7B2026.05 | 15 | — | — | |
| DocopilotSize=8B2026.05 | 12.1 | — | — | |
| R1-VL-7BModel Category=Reasoning Multimodal LLM, Parameters=7B2025.12 | 8.3 | — | — | |
| R1-VLSize=7B2026.05 | 8.3 | — | — | |
| BaseBase Model=Qwen3-VL-8B-Instruct2026.05 | — | 38.34 | — | |
| BaseBase Model=Qwen3-VL-2B-Instruct2026.05 | — | 26.87 | — | |
| BaseBase Model=Qwen2.5-VL-7B-Instruct2026.05 | — | 23.25 | — | |
| BaseBase Model=Qwen2-VL-7B-Instruct2026.05 | — | 11.75 | — | |
| Base (Qwen2-VL-8B Instruct)Temperature=0.62026.05 | — | 38.34 | — | |
| Claude-3.5Size=-2026.05 | — | — | 38.1 | |
| DAPOBase Model=Qwen2-VL-8B Instruct, Temperature=0.62026.05 | — | 38.75 | — | |
| DeepEyesSize=7B2026.05 | — | — | 18.1 | |
| DocopilotSize=8B2026.05 | — | — | 12.1 | |
| GPT-4oSize=-2026.05 | — | — | 32.7 | |
| GPT-4o-miniSize=-2026.05 | — | — | 27.3 | |
| GRPOBase Model=Qwen3-VL-8B-Instruct2026.05 | — | 41.34 | — | |
| GRPOBase Model=Qwen3-VL-2B-Instruct2026.05 | — | 28.5 | — | |
| Llava-OVSize=7B2026.05 | — | — | 18.3 | |
| MiMO-VL-InstructSize=7B2026.05 | — | — | 26.5 | |
| MINT-CoTBase Model=Qwen2-VL-7B-Instruct2026.05 | — | 19 | — | |
| MM-EurekaBase Model=Qwen2.5-VL-7B-Instruct2026.05 | — | 28.75 | — | |
| mPLUG-Owl3Size=7B2026.05 | — | — | 15 | |
| OpenVLThinkerSize=7B2026.05 | — | — | 26.6 | |
| PAPOBase Model=Qwen3-VL-8B-Instruct2026.05 | — | 39 | — | |
| PAPOBase Model=Qwen3-VL-2B-Instruct2026.05 | — | 27.25 | — | |
| PAPO_DBase Model=Qwen2-VL-8B Instruct, Temperature=0.62026.05 | — | 42.25 | — | |
| Pixel ReasonerSize=7B2026.05 | — | — | 19.8 | |
| Qwen2.5-VL-InstructSize=72B2026.05 | — | — | 38.5 | |
| Qwen2.5-VL-InstructSize=7B2026.05 | — | — | 21.5 | |
| R1-ShareVLBase Model=Qwen2.5-VL-7B-Instruct2026.05 | — | 30.5 | — | |
| R1-VLSize=7B2026.05 | — | — | 8.3 | |
| RAPOBase Model=Qwen3-VL-8B-Instruct2026.05 | — | 42.5 | — | |
| RAPOBase Model=Qwen3-VL-2B-Instruct2026.05 | — | 31.25 | — | |
| RAPO_DBase Model=Qwen2.5-VL-7B-Instruct2026.05 | — | 30.5 | — | |
| RAPO_DBase Model=Qwen2-VL-7B-Instruct2026.05 | — | 28.75 | — | |
| RAPO_DBase Model=Qwen2-VL-8B Instruct, Temperature=0.62026.05 | — | 45.5 | — | |
| RAPO_GBase Model=Qwen2.5-VL-7B-Instruct2026.05 | — | 29 | — | |
| RAPO_GBase Model=Qwen2-VL-7B-Instruct2026.05 | — | 25.25 | — | |
| TVCBase Model=Qwen2-VL-7B-Instruct2026.05 | — | 20.75 | — | |
| VL-RethinkerBase Model=Qwen2.5-VL-7B-Instruct2026.05 | — | 28.75 | — | |
| VL-RethinkerSize=7B2026.05 | — | — | 29.7 | |
| VL-ScalerSize=7B2026.05 | — | — | 31.3 | |
| VL-Scaler-MiMOSize=7B2026.05 | — | — | 32.1 |