Visual Mathematical Reasoning on MathVista
86.81AccuracyDynamo (Full)
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Dynamo (Full)Model=Doubao-Seed-2.0, Evolution Strategy=Full2026.06 | 86.81 | — | — | — | — | — | — | — | — | |
| Dynamo (Tool Only)Model=Doubao-Seed-2.0, Evolution Strategy=Tool Only2026.06 | 86.59 | — | — | — | — | — | — | — | — | |
| TVI-CoTBackbone=Qwen3-VL-32B2026.06 | 85.6 | — | — | — | — | — | — | — | — | |
| Dynamo (Skill Only)Model=Doubao-Seed-2.0, Evolution Strategy=Skill Only2026.06 | 85.56 | — | — | — | — | — | — | — | — | |
| Base AgentModel=Doubao-Seed-2.0, Evolution Strategy=None2026.06 | 85.44 | — | — | — | — | — | — | — | — | |
| InternVL3.5-8B-MastersSize=8B2025.12 | 85 | — | — | — | — | — | — | — | — | |
| MastersBase Model=InternVL3.5-8B2025.12 | 85 | — | — | — | — | — | — | — | — | |
| Supervised ETTCEnsemble Composition=Same Family Models2026.05 | 84.81 | — | — | — | — | — | — | — | 0.37 | |
| GLM-4.5V2025.12 | 84.6 | — | — | — | — | — | — | — | — | |
| ETTCEnsemble Composition=Same Family Models2026.05 | 84.44 | — | — | — | — | — | — | — | — | |
| Qwen3-VL-32B2025.12 | 83.8 | — | — | — | — | — | — | — | — | |
| Qwen3-VL-32BBackbone=Qwen3-VL-32B2026.06 | 83.8 | — | — | — | — | — | — | — | — | |
| SkillGraphModel=Qwen3-VL 8B-Instruct, Baseline=Complete2026.04 | 83.2 | — | — | — | — | — | — | — | — | |
| VotingEnsemble Composition=Same Family Models2026.05 | 83.15 | — | — | — | — | — | — | — | — | |
| SkillGraphModel=Qwen3-VL 8B-Instruct, Baseline=Random2026.04 | 82.8 | — | — | — | — | — | — | — | — | |
| Dynamo (Full)Model=Qwen3.5-27B, Evolution Strategy=Full2026.06 | 82.52 | — | — | — | — | — | — | — | — | |
| Dynamo (Tool Only)Model=Qwen3.5-27B, Evolution Strategy=Tool Only2026.06 | 82.41 | — | — | — | — | — | — | — | — | |
| InternVL3-8B-MastersSize=8B2025.12 | 82.3 | — | — | — | — | — | — | — | — | |
| MastersBase Model=InternVL3-8B2025.12 | 82.3 | — | — | — | — | — | — | — | — | |
| SkillGraphModel=Qwen3-VL 8B-Instruct, Baseline=Centralized2026.04 | 82.2 | — | — | — | — | — | — | — | — | |
| Octopus-8B (Ours)Rollout count (rollout.n)=8, Generation Time (Gen.)=344.7, Total training time per step=958.12026.02 | 82.1 | — | — | — | — | — | — | — | — | |
| GPT-52025.12 | 81.9 | — | — | — | — | — | — | — | — | |
| InternVL3.5-38B2025.12 | 81.9 | — | — | — | — | — | — | — | — | |
| MiMo-VL-7B-SFTLearning Protocol=Open-source Reasoning VLM2026.02 | 81.8 | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8B-MastersSize=8B2025.12 | 81.8 | — | — | — | — | — | — | — | — | |
| MastersBase Model=Qwen3-VL-8B2025.12 | 81.8 | — | — | — | — | — | — | — | — | |
| CompleteModel=Qwen3-VL 8B-Instruct2026.04 | 81.8 | — | — | — | — | — | — | — | — | |
| RandomModel=Qwen3-VL 8B-Instruct2026.04 | 81.6 | — | — | — | — | — | — | — | — | |
| Dynamo (Skill Only)Model=Qwen3.5-27B, Evolution Strategy=Skill Only2026.06 | 81.52 | — | — | — | — | — | — | — | — | |
| MiMo-VL-7B-RLLearning Protocol=Open-source Reasoning VLM2026.02 | 81.5 | — | — | — | — | — | — | — | — | |
| MiMo-VL-8BSize=8B2025.12 | 81.5 | — | — | — | — | — | — | — | — | |
| SkillGraphModel=Qwen3-VL 8B-Instruct, Baseline=Layered2026.04 | 81.5 | — | — | — | — | — | — | — | — | |
| Keye-VL-1.5-8BSize=8B2025.12 | 81.2 | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8B-Instruct + GSPORollout count (rollout.n)=16, Generation Time (Gen.)=688.9, Total training time per step=1322.8, Learning Protocol=RLVR2026.02 | 81 | — | — | — | — | — | — | — | — | |
| Gemini-2.5-Pro2025.12 | 80.9 | — | — | — | — | — | — | — | — | |
| Base AgentModel=Qwen3.5-27B, Evolution Strategy=None2026.06 | 80.88 | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8B-Instruct + GRPORollout count (rollout.n)=16, Generation Time (Gen.)=679.1, Total training time per step=1428.4, Learning Protocol=RLVR2026.02 | 80.8 | — | — | — | — | — | — | — | — | |
| SkillGraphModel=Qwen3-VL 8B-Instruct, Baseline=Linear2026.04 | 80.8 | — | — | — | — | — | — | — | — | |
| GLM-4.1V-9BSize=9B2025.12 | 80.7 | — | — | — | — | — | — | — | — | |
| Keye-VL-8BSize=8B2025.12 | 80.7 | — | — | — | — | — | — | — | — | |
| CentralizedModel=Qwen3-VL 8B-Instruct2026.04 | 80.5 | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8B-Instruct + DAPORollout count (rollout.n)=16, Learning Protocol=RLVR2026.02 | 80.3 | — | — | — | — | — | — | — | — | |
| LayeredModel=Qwen3-VL 8B-Instruct2026.04 | 80.2 | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8B-ThinkingLearning Protocol=Open-source Reasoning VLM2026.02 | 79.8 | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8B-Instruct + SRPORollout count (rollout.n)=16, Generation Time (Gen.)=816.6, Total training time per step=1543.3, Learning Protocol=RLVR2026.02 | 79.8 | — | — | — | — | — | — | — | — | |
| LinearModel=Qwen3-VL 8B-Instruct2026.04 | 79.7 | — | — | — | — | — | — | — | — | |
| Supervised ETTCEnsemble Composition=Similar Size Models2026.05 | 79.63 | — | — | — | — | — | — | — | 3.7 | |
| Qwen3-VL-4B-MastersSize=4B2025.12 | 79.6 | — | — | — | — | — | — | — | — | |
| Doubao-pro-1.5Version=1.5 Pro2026.01 | 79.5 | — | 77.7 | 88.9 | 86 | 82.3 | 62 | — | — | |
| Qwen3-VL-8B-Instruct + GRPORollout count (rollout.n)=8, Generation Time (Gen.)=364.9, Total training time per step=845.1, Learning Protocol=RLVR2026.02 | 79.1 | — | — | — | — | — | — | — | — | |
| GPT-5-Mini2025.12 | 79.1 | — | — | — | — | — | — | — | — | |
| InternVL3-78B2025.12 | 79 | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-7B-MastersSize=7B2025.12 | 78.8 | — | — | — | — | — | — | — | — | |
| InternVL3.5-4B-MastersSize=4B2025.12 | 78.8 | — | — | — | — | — | — | — | — | |
| MastersBase Model=Qwen2.5-VL-7B2025.12 | 78.8 | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8B-Instruct + GSPORollout count (rollout.n)=8, Generation Time (Gen.)=361.7, Total training time per step=753.1, Learning Protocol=RLVR2026.02 | 78.5 | — | — | — | — | — | — | — | — | |
| InternVL3.5-8BSize=8B2025.12 | 78.4 | — | — | — | — | — | — | — | — | |
| FULLModel=Qwen3-VL (4B), Selection Ratio=100%2026.05 | 78.4 | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-32B-Instruct + NoisyRolloutStudent Model=Qwen2.5-VL-32B-Instruct, Distillation/Training Method=NoisyRollout2026.03 | 78.3 | — | — | — | — | — | — | — | — | |
| Dynamo (Full)Model=GPT-5.4, Evolution Strategy=Full2026.06 | 78.15 | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8B-Instruct + SRPORollout count (rollout.n)=8, Generation Time (Gen.)=410.9, Total training time per step=895.8, Learning Protocol=RLVR2026.02 | 77.9 | — | — | — | — | — | — | — | — | |
| SkillGraphModel=InternVL3-8B, Baseline=Complete2026.04 | 77.8 | — | — | — | — | — | — | — | — | |
| SkillGraphModel=InternVL3-8B, Baseline=Random2026.04 | 77.6 | — | — | — | — | — | — | — | — | |
| Dynamo (Skill Only)Model=GPT-5.4, Evolution Strategy=Skill Only2026.06 | 77.52 | — | — | — | — | — | — | — | — | |
| InternVL3.5-2B-MastersSize=2B2025.12 | 77.5 | — | — | — | — | — | — | — | — | |
| DirectAnswerModel=Qwen3-VL 8B-Instruct2026.04 | 77.3 | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8BSize=8B2025.12 | 77.2 | — | — | — | — | — | — | — | — | |
| InternVL3.5-4BSize=4B2025.12 | 77.1 | — | — | — | — | — | — | — | — | |
| PRCO-7BBackbone=Qwen2.5-VL-7B2026.03 | 77.1 | — | — | — | — | — | — | — | — | |
| Shuffle-R1-Qwen-7BTraining Strategy=Zero RL2025.08 | 77 | — | — | — | — | — | — | — | — | |
| Qwen3-VL-8B-InstructLearning Protocol=Base VLM2026.02 | 76.9 | — | — | — | — | — | — | — | — | |
| COGFLOWParameter Scale=7B2026.01 | 76.8 | — | 70.4 | 93.1 | 73.7 | 86.9 | 59.3 | — | — | |
| PAPO-D-7BBackbone=Qwen2.5-VL-7B2026.03 | 76.7 | — | — | — | — | — | — | — | — | |
| Dynamo (Tool Only)Model=GPT-5.4, Evolution Strategy=Tool Only2026.06 | 76.63 | — | — | — | — | — | — | — | — | |
| VPPO-7BBackbone=Qwen2.5-VL-7B2026.03 | 76.6 | — | — | — | — | — | — | — | — | |
| SkillGraphModel=InternVL3-8B, Baseline=Centralized2026.04 | 76.5 | — | — | — | — | — | — | — | — | |
| SkillGraphModel=InternVL3-8B, Baseline=Layered2026.04 | 76.4 | — | — | — | — | — | — | — | — | |
| OPD T→VTraining Strategy=Static On-policy Policy Distillation (Text to Image)2026.04 | 76.05 | — | — | — | — | — | — | — | — | |
| ETTCEnsemble Composition=Similar Size Models2026.05 | 75.93 | — | — | — | — | — | — | — | — | |
| CodePercept-32B-S1LLM Solver=Qwen3-235A22-Thinking [53]2026.03 | 75.9 | — | — | — | — | — | — | — | — | |
| SkillGraphModel=InternVL3-8B, Baseline=Linear2026.04 | 75.9 | — | — | — | — | — | — | — | — | |
| Base AgentModel=GPT-5.4, Evolution Strategy=None2026.06 | 75.89 | — | — | — | — | — | — | — | — | |
| CoPDTraining Strategy=Co-Evolving Policy Distillation2026.04 | 75.75 | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-3B-MastersSize=3B2025.12 | 75.6 | — | — | — | — | — | — | — | — | |
| DualMindVLMSize=7B, Strategy=RL2025.11 | 75.6 | — | — | — | — | — | — | 184 | — | |
| Dynamo (Full)Model=o4-mini, Evolution Strategy=Full2026.06 | 75.59 | — | — | — | — | — | — | — | — | |
| GRPOBackbone=Qwen2.5-VL-7B2026.03 | 75.4 | — | — | — | — | — | — | — | — | |
| Base AgentModel=o4-mini, Evolution Strategy=None2026.06 | 75.33 | — | — | — | — | — | — | — | — | |
| CompleteModel=InternVL3-8B2026.04 | 75.3 | — | — | — | — | — | — | — | — | |
| RandomModel=InternVL3-8B2026.04 | 75.2 | — | — | — | — | — | — | — | — | |
| ThinkLiteSize=7B, Strategy=RL2025.11 | 75.1 | — | — | — | — | — | — | 247 | — | |
| Image-ExpertTraining Strategy=Image-Specific Expert2026.04 | 75.1 | — | — | — | — | — | — | — | — | |
| Mixed RLVRTraining Strategy=Joint optimization on combined budget2026.04 | 75.1 | — | — | — | — | — | — | — | — | |
| VL-RethinkerSize=7B, Strategy=RL2025.11 | 74.9 | — | — | — | — | — | — | 268 | — | |
| VL-RethinkerSize=7B2026.05 | 74.9 | — | — | — | — | — | — | — | — | |
| Gemini2.5-ProLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 74.8 | — | — | — | — | — | — | — | — | |
| DAPOBackbone=Qwen2.5-VL-7B2026.03 | 74.8 | — | — | — | — | — | — | — | — | |
| Qwen2.5-VL-InstructSize=72B2026.05 | 74.8 | — | — | — | — | — | — | — | — | |
| QvQ-72B-PreviewModel Access=Open-source, Model Scale=72B2026.06 | 74.8 | — | — | — | — | — | — | — | — | |
| InternVL2.5-38BModel Access=Open-source, Model Scale=38B2026.06 | 74.7 | — | — | — | — | — | — | — | — |