Visual Mathematical Reasoning on MathVerse
78.8AccuracyTVI-CoT
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| TVI-CoTBackbone=Qwen3-VL-32B2026.06 | 78.8 | — | — | — | — | |
| Qwen3-VL-32BBackbone=Qwen3-VL-32B2026.06 | 76.8 | — | — | — | — | |
| CodePercept-32B-S1LLM Solver=Qwen3-235A22-Thinking [53]2026.03 | 73.56 | — | — | — | — | |
| Gemini2.5-ProLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 73.47 | — | — | — | — | |
| CodePercept-32B-S1LLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 71.7 | — | — | — | — | |
| Qwen3-VL-32B-InstructLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 71.09 | — | — | — | — | |
| PDCRBackbone=Qwen3-VL-8B-Instruct2026.05 | 70.6 | — | — | — | — | |
| Qwen3-VL-235A22B-InstructLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 70.08 | — | — | — | — | |
| Qwen3-VL-32B-InstructLLM Solver=Qwen3-235A22-Thinking [53]2026.03 | 69.9 | — | — | — | — | |
| PACRBackbone=Qwen3-VL-8B-Instruct2026.05 | 69.9 | — | — | — | — | |
| GPT5-ThinkingLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 69.56 | — | — | — | — | |
| GRPOBackbone=Qwen3-VL-8B-Instruct2026.05 | 69.3 | — | — | — | — | |
| Qwen3-VL-8B-ThinkingLearning Protocol=Open-source Reasoning VLM2026.02 | 69.2 | — | — | — | — | |
| DAPOBackbone=Qwen3-VL-8B-Instruct2026.05 | 69.2 | — | — | — | — | |
| Qwen3-VL-8B-Instruct + DAPORollout count (rollout.n)=16, Learning Protocol=RLVR2026.02 | 68.5 | — | — | — | — | |
| Octopus-8B (Ours)Rollout count (rollout.n)=8, Generation Time (Gen.)=344.7, Total training time per step=958.12026.02 | 68.5 | — | — | — | — | |
| Qwen3-VL-8B-Instruct + GSPORollout count (rollout.n)=16, Generation Time (Gen.)=688.9, Total training time per step=1322.8, Learning Protocol=RLVR2026.02 | 68.4 | — | — | — | — | |
| Zero-shot InferenceBackbone=Qwen3-VL-8B-Instruct2026.05 | 68.1 | — | — | — | — | |
| CodePercept-8B-S1LLM Solver=Qwen3-235A22-Thinking [53]2026.03 | 67.95 | — | — | — | — | |
| Gemini-2.0 proTraining Strategy=Close-source2025.08 | 67.3 | — | — | — | — | |
| CodePercept-4B-S1LLM Solver=Qwen3-235A22-Thinking [53]2026.03 | 66.73 | — | — | — | — | |
| CodePercept-8B-S1LLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 66.52 | — | — | — | — | |
| Qwen3-VL-30A3B-InstructLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 66.44 | — | — | — | — | |
| Qwen3-VL-4B-InstructLLM Solver=Qwen3-235A22-Thinking [53]2026.03 | 66.39 | — | — | — | — | |
| Qwen3-VL-8B-Instruct + GRPORollout count (rollout.n)=16, Generation Time (Gen.)=679.1, Total training time per step=1428.4, Learning Protocol=RLVR2026.02 | 66.3 | — | — | — | — | |
| Qwen3-VL-8B-Instruct + GRPORollout count (rollout.n)=8, Generation Time (Gen.)=364.9, Total training time per step=845.1, Learning Protocol=RLVR2026.02 | 66 | — | — | — | — | |
| Claude4-Sonnet2026.02 | 65.9 | — | — | — | 64.1 | |
| CodePercept-4B-S1LLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 65.59 | — | — | — | — | |
| Qwen3-VL-4B-InstructLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 64.59 | — | — | — | — | |
| Qwen3-VL-8B-Instruct + SRPORollout count (rollout.n)=16, Generation Time (Gen.)=816.6, Total training time per step=1543.3, Learning Protocol=RLVR2026.02 | 64.2 | — | — | — | — | |
| Qwen3-VL-8B-Instruct + SRPORollout count (rollout.n)=8, Generation Time (Gen.)=410.9, Total training time per step=895.8, Learning Protocol=RLVR2026.02 | 64.1 | — | — | — | — | |
| Qwen3-VL-8B-InstructLLM Solver=Qwen3-235A22-Thinking [53]2026.03 | 63.88 | — | — | — | — | |
| Qwen3-VL-8B-InstructLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 63.75 | — | — | — | — | |
| Qwen3-VL-8B-Instruct + GSPORollout count (rollout.n)=8, Generation Time (Gen.)=361.7, Total training time per step=753.1, Learning Protocol=RLVR2026.02 | 63.7 | — | — | — | — | |
| InternVL3.5-8B-RLLearning Protocol=Open-source Reasoning VLM2026.02 | 61.5 | — | — | — | — | |
| MiMo-VL-7B-SFTLearning Protocol=Open-source Reasoning VLM2026.02 | 61.4 | — | — | — | — | |
| MiMo-VL-7B-RLLearning Protocol=Open-source Reasoning VLM2026.02 | 61 | — | — | — | — | |
| Gemini-2.0-FlashModel Access=Closed-source2026.06 | 59.3 | — | — | — | — | |
| Qwen2.5-VL-32B-Instruct + NoisyRolloutStudent Model=Qwen2.5-VL-32B-Instruct, Distillation/Training Method=NoisyRollout2026.03 | 58.86 | — | — | — | — | |
| MiniCPM-V-4.5LLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 57.84 | — | — | — | — | |
| Qwen2.5-VL-72BParameters=72B2026.02 | 57.6 | — | — | — | 61.9 | |
| QvQ-72B-PreviewModel Access=Open-source, Model Scale=72B2026.06 | 57.6 | — | — | — | — | |
| DPEBackbone=Qwen3-VL-8B-Instruct, Iteration=32026.02 | 57.18 | — | — | — | 64.39 | |
| OpenAI-o1Learning Protocol=Closed-source VLM2026.02 | 57 | — | — | — | — | |
| o1Training Strategy=Close-source2025.08 | 57 | — | — | — | — | |
| o1Model Access=Closed-source2026.06 | 57 | — | — | — | — | |
| Claude-Opus 4.1-ThinkingLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 56.19 | — | — | — | — | |
| Qwen2.5-VL-72BLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 55.4 | — | — | — | — | |
| GLM-4.1V-9BLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 54.47 | — | — | — | — | |
| VL-RethinkerModel Scale=7B2025.09 | 54.2 | — | — | — | — | |
| Shuffle-R1-Qwen-7BTraining Strategy=Zero RL2025.08 | 53.9 | — | — | — | — | |
| InternVL3.5-8BLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 53.4 | — | — | — | — | |
| VAPOModel Scale=7B2025.09 | 53.3 | — | — | — | — | |
| NoisyRolloutTeacher Model=Qwen2.5-VL-32B-Instruct + NoisyRollout, Student Model=Qwen2.5-VL-7B-Instruct, Distillation/Training Method=NoisyRollout2026.03 | 53.14 | — | — | — | — | |
| DeepEyesV22026.02 | 52.9 | — | — | — | — | |
| Qwen3-VL-8B-InstructLearning Protocol=Base VLM2026.02 | 52.6 | — | — | — | — | |
| Claude-3.7-SonnetLearning Protocol=Closed-source VLM2026.02 | 52 | — | — | — | — | |
| Claude-3.7-SonnetTraining Strategy=Close-source2025.08 | 52 | — | — | — | — | |
| Claude-3.7-SonnetModel Access=Closed-source2026.06 | 52 | — | — | — | — | |
| Intern-S1-8BLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 51.9 | — | — | — | — | |
| FULLModel=Qwen3-VL (4B), Selection Ratio=100%2026.05 | 51.8 | — | — | — | — | |
| VL-Rethinker-7BTraining Strategy=Zero RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 51.7 | — | — | — | — | |
| InternVL2.5-78BModel Access=Open-source, Model Scale=78B2026.06 | 51.7 | — | — | — | — | |
| REOPOLDTeacher Model=Qwen2.5-VL-32B-Instruct + NoisyRollout, Student Model=Qwen2.5-VL-7B-Instruct, Distillation/Training Method=REOPOLD2026.03 | 51.43 | — | — | — | — | |
| SaEIFinetuned=true, Group Size=122025.12 | 51.06 | — | — | — | — | |
| KL-CovFinetuned=true, Group Size=122025.12 | 50.97 | — | — | — | — | |
| Vanilla GRPOFinetuned=true, Group Size=122025.12 | 50.95 | — | — | — | — | |
| GPT-4oTraining Strategy=Close-source2025.08 | 50.8 | — | — | — | — | |
| GRPOTeacher Model=Qwen2.5-VL-32B-Instruct + NoisyRollout, Student Model=Qwen2.5-VL-7B-Instruct, Distillation/Training Method=GRPO2026.03 | 50.8 | — | — | — | — | |
| GPT-4oModel Access=Closed-source2026.06 | 50.8 | — | — | — | — | |
| PAPOModel Scale=7B, Decoding Strategy=greedy decoding2025.09 | 50.4 | — | — | — | — | |
| MM-EurekaModel Scale=7B2025.09 | 50.3 | — | — | — | — | |
| GPT-4oLearning Protocol=Closed-source VLM2026.02 | 50.2 | — | — | — | — | |
| GPT-4o2026.02 | 50.2 | — | — | — | 56.1 | |
| NoisyRollout-7B-K12Training Strategy=Zero RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 50.1 | — | — | — | — | |
| KeyeVL1.5-8BLLM Solver=Qwen3-30A3-Thinking [53]2026.03 | 49.95 | — | — | — | — | |
| InternVL2.5-38BModel Access=Open-source, Model Scale=38B2026.06 | 49.9 | — | — | — | — | |
| MM-Eureka-Qwen-7BTraining Strategy=Zero RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 49.6 | — | — | — | — | |
| PRCO-7BBackbone=Qwen2.5-VL-7B2026.03 | 49.49 | — | — | — | — | |
| RANDOMModel=Qwen3-VL (4B), Selection Ratio=20%2026.05 | 49.4 | — | — | — | — | |
| Qwen2.5-VL-72BModel Access=Open-source, Model Scale=72B2026.06 | 49.4 | — | — | — | — | |
| NoisyRollout-7BModel Access=Open-source, Model Scale=7B, Learning Strategy=Visual-focused RL Fine-tuning2026.06 | 49.16 | — | — | — | — | |
| Qwen2.5-VL-7B-Instruct + VEPOModel Access=Open-source, Model Scale=7B, Learning Strategy=VEPO2026.06 | 48.93 | — | — | — | — | |
| VLAA-Thinker-7BTraining Strategy=Cold-Start + RL, Evaluation Toolkit=vLLM with custom scripts2025.08 | 48.9 | — | — | — | — | |
| Qwen2.5-VL-7B-Instruct + GRPO (Top 40% high entropy)Model Access=Open-source, Model Scale=7B, Learning Strategy=GRPO, Entropy Selection Strategy=Top 40% high entropy2026.06 | 48.86 | — | — | — | — | |
| DAPOBackbone=Qwen2.5-VL-7B2026.03 | 48.73 | — | — | — | — | |
| Qwen2.5-VL-7B-Instruct + GRPOModel Access=Open-source, Model Scale=7B, Learning Strategy=GRPO2026.06 | 48.53 | — | — | — | — | |
| Honeybee-Remake-SEED-200KModel=Qwen3-VL (4B), Selection Ratio=20%2026.05 | 48.4 | — | — | — | — | |
| Qwen2.5-VL-7B-Instruct + GRPO (Top 20% high entropy)Model Access=Open-source, Model Scale=7B, Learning Strategy=GRPO, Entropy Selection Strategy=Top 20% high entropy2026.06 | 48.25 | — | — | — | — | |
| R1-ShareVL-7BBackbone=Qwen2.5-VL-7B2026.03 | 48.22 | — | — | — | — | |
| VPPO-7BModel Access=Open-source, Model Scale=7B, Learning Strategy=Visual-focused RL Fine-tuning2026.06 | 48.21 | — | — | — | — | |
| Qwen2.5-VL-32BModel Access=Open-source, Model Scale=32B2026.06 | 48.2 | — | — | — | — | |
| Active-ZeroModel Backbone=Qwen2.5-VL-7B-Instruct2026.02 | 48.16 | — | — | — | — | |
| PAPO-DAPO-7BModel Access=Open-source, Model Scale=7B, Learning Strategy=Visual-focused RL Fine-tuning2026.06 | 48.07 | — | — | — | — | |
| NoisyRolloutFinetuned=true, Group Size=122025.12 | 47.74 | — | — | — | — | |
| RKLTeacher Model=Qwen2.5-VL-32B-Instruct + NoisyRollout, Student Model=Qwen2.5-VL-7B-Instruct, Distillation/Training Method=RKL2026.03 | 47.71 | — | — | — | — | |
| R1-ShareVLModel Access=Open-source, Model Scale=7B, Learning Strategy=Visual-focused RL Fine-tuning2026.06 | 47.39 | — | — | — | — | |
| PAPO-D-7BBackbone=Qwen2.5-VL-7B2026.03 | 47.33 | — | — | — | — | |
| DeepEyes2026.02 | 47.3 | — | — | — | — | |
| VPPO-7BBackbone=Qwen2.5-VL-7B2026.03 | 47.2 | — | — | — | — |