Multimodal Mathematical Reasoning on DynaMath (accuracy)
69.2Accuracy (DynaMath)Metis
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MetisModel Category=Agentic Multimodal Models2026.04 | 69.2 | — | |
| Qwen3-VL-8B-InstructModel Category=Open-source Models2026.04 | 65.5 | — | |
| CEPOBackbone=Qwen3-VL-4B-Instruct2026.05 | 65.37 | — | |
| RLSDBackbone=Qwen3-VL-4B-Instruct2026.05 | 65.07 | — | |
| BaseBackbone=Qwen3-VL-4B-Instruct2026.05 | 64.59 | — | |
| GRPOBackbone=Qwen3-VL-4B-Instruct2026.05 | 63.97 | — | |
| GPT-4oTool=Code, Param Size=-2025.11 | 61.9 | — | |
| OPSDBackbone=Qwen3-VL-4B-Instruct2026.05 | 61.8 | — | |
| SDPOBackbone=Qwen3-VL-4B-Instruct2026.05 | 61.58 | — | |
| DeepEyesV2Tool=General, Param Size=7B2025.11 | 57.2 | — | |
| DeepEyesV2Model Category=Agentic Multimodal Models2026.04 | 57.2 | — | |
| DeepEyesTool=Crop, Param Size=7B2025.11 | 55 | — | |
| DeepEyesModel Category=Agentic Multimodal Models2026.04 | 55 | — | |
| Qwen-2.5-VLTool=✗, Param Size=7B2025.11 | 53.3 | — | |
| Qwen-2.5-VL-7B-InstructModel Category=Open-source Models2026.04 | 53.3 | — | |
| CEPOBackbone=Qwen3-VL-2B-Instruct2026.05 | 51.44 | — | |
| VOLDImages in FT=false2025.10 | 50.7 | — | |
| GRPOBackbone=Qwen3-VL-2B-Instruct2026.05 | 50.36 | — | |
| RLSDBackbone=Qwen3-VL-2B-Instruct2026.05 | 50.36 | — | |
| BaseBackbone=Qwen3-VL-2B-Instruct2026.05 | 50.08 | — | |
| VLAA-Thinker 3BImages in FT=true2025.10 | 47.5 | — | |
| XReasoner-3B (repl.)Images in FT=false2025.10 | 47.2 | — | |
| OPSDBackbone=Qwen3-VL-2B-Instruct2026.05 | 46.85 | — | |
| SDPOBackbone=Qwen3-VL-2B-Instruct2026.05 | 46.65 | — | |
| Qwen2.5-VL-3BImages in FT=-2025.10 | 42.7 | — | |
| VLM-R1 3B-MathImages in FT=true2025.10 | 42.7 | — | |
| Vision-R1Training Data=Mixed Data, Backbone=Qwen2.5-VL-7B2026.03 | 25.2 | — | |
| SelfJudgeTraining Data=Geo3K, Backbone=Qwen2.5-VL-7B2026.03 | 24.2 | — | |
| RL(GRPO)Training Data=Geo3K, Backbone=Qwen2.5-VL-7B2026.03 | 23.8 | — | |
| RL(GRPO)Training Data=GeoQA, Backbone=Qwen2.5-VL-7B2026.03 | 23.4 | — | |
| RL(GRPO)Training Data=MMR1, Backbone=Qwen2.5-VL-7B2026.03 | 23.3 | — | |
| SelfJudgeTraining Data=GeoQA, Backbone=Qwen2.5-VL-7B2026.03 | 23.2 | — | |
| SelfJudgeTraining Data=MMR1, Backbone=Qwen2.5-VL-7B2026.03 | 23 | — | |
| SFTTraining Data=MMR1, Backbone=Qwen2.5-VL-7B2026.03 | 22.9 | — | |
| SFTTraining Data=GeoQA, Backbone=Qwen2.5-VL-7B2026.03 | 22.6 | — | |
| SFTTraining Data=Geo3K, Backbone=Qwen2.5-VL-7B2026.03 | 22.1 | — | |
| MM-UPTTraining Data=GeoQA, Backbone=Qwen2.5-VL-7B2026.03 | 21.9 | — | |
| R1-Onevision-7BTraining Data=Mixed Data, Backbone=Qwen2.5-VL-7B2026.03 | 21.8 | — | |
| MM-UPTTraining Data=MMR1, Backbone=Qwen2.5-VL-7B2026.03 | 21.8 | — | |
| VisionZeroTraining Data=CLEVR, Backbone=Qwen2.5-VL-7B2026.03 | 21.7 | — | |
| MM-UPTTraining Data=Geo3K, Backbone=Qwen2.5-VL-7B2026.03 | 21.4 | — | |
| VisionZeroTraining Data=ImgEdit, Backbone=Qwen2.5-VL-7B2026.03 | 21.3 | — | |
| OpenVLThinker-8BTraining Data=Mixed Data, Backbone=Qwen2.5-VL-7B2026.03 | 21.2 | — | |
| EvoLMMTraining Data=Multi-Bench, Backbone=Qwen2.5-VL-7B2026.03 | 21 | — | |
| Qwen2.5-VL-7BBackbone=Qwen2.5-VL-7B2026.03 | 20.3 | — | |
| Claude-3.5-SonnetModel Source Category=Closed-Source Models2026.04 | — | 36.6 | |
| Gemini-1.5-ProModel Source Category=Closed-Source Models2026.04 | — | 40.7 | |
| GPT-4o-20240513Model Source Category=Closed-Source Models2026.04 | — | 34.5 | |
| InternVL2.5-26BModel Source Category=Open-Source Models2026.04 | — | 11.4 | |
| InternVL2.5-8BModel Source Category=Open-Source Models2026.04 | — | 9.4 | |
| InternVL3-9BModel Source Category=Open-Source Models2026.04 | — | 26.7 | |
| LLaVA-OneVision-72BModel Source Category=Open-Source Models2026.04 | — | 15.6 | |
| MiniCPM-V2.6Model Source Category=Open-Source Models2026.04 | — | 9.8 | |
| Qwen2-VL-7BModel Source Category=Open-Source Models2026.04 | — | 18.1 | |
| Qwen2.5-VL-7B-InstructModel Source Category=Base Model2026.04 | — | 21 | |
| Qwen2.5-VL-7B-InstructModel Source Category=Base Model, Chain-of-Thought (CoT)=true2026.04 | — | 23.7 | |
| Qwen2.5-VL-7B-InstructModel Source Category=Ours2026.04 | — | 21 | |
| Qwen2.5-VL-7B-Instruct + cold startModel Source Category=Ours, Training Strategy=cold start2026.04 | — | 26.8 | |
| Qwen2.5-VL-7B-Instruct + RFTModel Source Category=Ours, Training Strategy=RFT2026.04 | — | 29.2 |