Multimodal Reasoning on WeMath
78AccuracyGemini-2.5-Pro
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| Gemini-2.5-ProModel Category=Closed-source models2026.06 | 78 | — | — | — | — | |
| GPT-5Model Category=Closed-source models2026.06 | 77.1 | — | — | — | — | |
| TGRL-DAPOData=39K, Base Model=Qwen-2.5-VL-7B2026.03 | 72.2 | — | — | — | — | |
| Perception-R1Data=1.4K2026.03 | 72 | — | — | — | — | |
| TGRL-DAPOData=2.1K, Base Model=Qwen-2.5-VL-7B2026.03 | 71.05 | — | — | — | — | |
| Groupwise Ranking RewardReward Method=Groupwise Ranking Reward, Training Budget=Matched, Training Epochs=22026.04 | 70.8 | — | — | — | — | |
| NoisyRolloutData=6.4K2026.03 | 70.3 | — | — | — | — | |
| TGRL-GRPOData=2.1K, Base Model=Qwen-2.5-VL-7B2026.03 | 70.29 | — | — | — | — | |
| TGRL-GRPOData=39K, Base Model=Qwen-2.5-VL-7B2026.03 | 70.23 | — | — | — | — | |
| DAPOData=2.1K, Base Model=Qwen-2.5-VL-7B2026.03 | 69.3 | — | — | — | — | |
| ThinkLiteData=11K2026.03 | 69.2 | — | — | — | — | |
| DAPOData=39K, Base Model=Qwen-2.5-VL-7B2026.03 | 68.84 | — | — | — | — | |
| CFPOGBackbone=Qwen2.5-VL-7B, Hyperparameters=gamma=0.02 + No Ent2026.06 | 68.79 | — | — | — | — | |
| GRPOData=39K, Base Model=Qwen-2.5-VL-7B2026.03 | 68.2 | — | — | — | — | |
| GRPOBackbone=Qwen2.5-VL-7B2026.06 | 68.12 | — | — | — | — | |
| VLAA-ThinkerData=25K2026.03 | 67.9 | — | — | — | — | |
| OpenVLThinkerData=50K2026.03 | 67.8 | — | — | — | — | |
| GRPOData=2.1K, Base Model=Qwen-2.5-VL-7B2026.03 | 67.4 | — | — | — | — | |
| NPOstage=early + late-stage2026.04 | 66.95 | — | — | — | — | |
| PAPOGBackbone=Qwen2.5-VL-7B2026.06 | 66.55 | — | — | — | — | |
| AutoNPO2026.04 | 66 | — | — | — | — | |
| No CompressionBackbone=Qwen3-VL, Evaluation Protocol=SFT2026.02 | 63.8 | 562.6 | 8.8 | — | — | |
| Pointwise GRReward Method=Pointwise GR, Training Budget=Matched, Training Epochs=22026.04 | 63.8 | — | — | — | — | |
| Prune-on-LogicBackbone=Qwen3-VL, Evaluation Protocol=SFT2026.02 | 63.4 | 419.5 | 6.6 | — | — | |
| Qwen-2.5-VL-7BData=-2026.03 | 63.1 | — | — | — | — | |
| XMCCBackbone=Qwen3-VL, Evaluation Protocol=SFT2026.02 | 63 | 97.7 | 1.6 | — | — | |
| NPOstage=early-stage only2026.04 | 62.76 | — | — | — | — | |
| ExGRPOtype=historical replay2026.04 | 62.67 | — | — | — | — | |
| RLEPtype=far future2026.04 | 62.48 | — | — | — | — | |
| SFT-MModel=Qwen2.5-VL-7B, Paradigm=SFT-M2026.02 | 62.3 | — | — | — | — | |
| R1-OnevisionData=165K2026.03 | 62.1 | — | — | — | — | |
| StepEntropyBackbone=Qwen3-VL, Evaluation Protocol=SFT2026.02 | 61.6 | 313.5 | 5.1 | — | — | |
| Socratic-Solver-Geo (Stage3)Data Scale=2.5k, Curriculum Stage=32026.02 | 61.58 | — | — | — | — | |
| PRMReward Method=PRM, Training Budget=Matched, Training Epochs=22026.04 | 60.8 | — | — | — | — | |
| SFTModel=Qwen2.5-VL-7B, Paradigm=SFT2026.02 | 60.23 | — | — | — | — | |
| SFT-RSModel=Qwen2.5-VL-7B, Paradigm=SFT-RS2026.02 | 59.2 | — | — | — | — | |
| RLVRReference Checkpoint=RLVR2026.04 | 58.8 | — | — | — | — | |
| Athena-PRMPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 58.7 | — | — | — | — | |
| TrustGeoGenData Scale=10k2026.02 | 58.61 | — | — | — | — | |
| PGPS9KData Scale=10k2026.02 | 58.53 | — | — | — | — | |
| Socratic-Solver-Geo (Stage2)Data Scale=1k, Curriculum Stage=22026.02 | 58.21 | — | — | — | — | |
| KD (Our Synthesis)Data Scale=2.5k2026.02 | 58.15 | — | — | — | — | |
| KD (Geo3K)Data Scale=3k2026.02 | 58.02 | — | — | — | — | |
| RLSDBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 58 | — | — | — | — | |
| GRPOModel=Qwen2.5-VL-7B, Paradigm=GRPO2026.02 | 57.99 | — | — | — | — | |
| GeoReasoningData Scale=10k2026.02 | 57.9 | — | — | — | — | |
| Qwen2.5-VL-7B-InstructMode=Zero-shot2026.02 | 57.59 | — | — | — | — | |
| R-CoTData Scale=7.2k2026.02 | 57.59 | — | — | — | — | |
| Socratic-Solver-Geo (Stage1)Data Scale=0.4k, Curriculum Stage=12026.02 | 57.54 | — | — | — | — | |
| Geo170k (G-LLaVA)Data Scale=10k2026.02 | 57.44 | — | — | — | — | |
| Qwen2.5-VL-7B-ITReference Checkpoint=Qwen2.5-VL-7B-IT2026.04 | 56.6 | — | — | — | — | |
| GRPOBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 56.57 | — | — | — | — | |
| GRPOtype=pure on-policy2026.04 | 56.57 | — | — | — | — | |
| No CompressionBackbone=Qwen2.5-VL, Evaluation Protocol=SFT2026.02 | 55.8 | 572.5 | 10.3 | — | — | |
| Athena-ORMPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 55.6 | — | — | — | — | |
| GPT-4.1VLM Category=Proprietary VLMs2025.09 | 55.5 | — | — | — | — | |
| Prune-on-LogicBackbone=Qwen2.5-VL, Evaluation Protocol=SFT2026.02 | 55.4 | 460.5 | 8.3 | — | — | |
| OPSDBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 54.95 | — | — | — | — | |
| XMCCBackbone=Qwen2.5-VL, Evaluation Protocol=SFT2026.02 | 54.9 | 99.6 | 1.8 | — | — | |
| SFT-MModel=Qwen2.5-VL-3B, Paradigm=SFT-M2026.02 | 54.89 | — | — | — | — | |
| Self-consistencyPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 54.8 | — | — | — | — | |
| GRPO + OPSDBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 54.76 | — | — | — | — | |
| StepEntropyBackbone=Qwen2.5-VL, Evaluation Protocol=SFT2026.02 | 54.6 | 334.2 | 6.1 | — | — | |
| SFTModel=Qwen2.5-VL-3B, Paradigm=SFT2026.02 | 54.14 | — | — | — | — | |
| Base LLMBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 54.1 | — | — | — | — | |
| Qwen3-VL-8B-Instruct2026.04 | 54.1 | — | — | — | — | |
| TACO (Ours-7B)Model Category=Code-tool / visual-agent models2026.06 | 53.1 | — | — | — | — | |
| GRPOModel=Qwen2.5-VL-3B, Paradigm=GRPO2026.02 | 52.82 | — | — | — | — | |
| SFT-RSModel=Qwen2.5-VL-3B, Paradigm=SFT-RS2026.02 | 52.41 | — | — | — | — | |
| LUFFYtype=external teacher2026.04 | 52.38 | — | — | — | — | |
| SDPOBase Model=Qwen3-VL-8B-Instruct, Max Context Length=81922026.04 | 52.19 | — | — | — | — | |
| Qwen2.5-VL-72B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=GPT-OSS-120B-A5B2025.06 | 52.1 | — | — | — | — | |
| MathCoder-VL-8BModel Category=Code-tool / visual-agent models2026.06 | 52.1 | — | — | — | — | |
| Ovis2-34BModel Category=Open-Source2025.06 | 51.9 | — | — | — | — | |
| InternVL2.5-38B + VRPRM w/o RLBase Model=InternVL2.5-38B, Selection Strategy=Best-of-8, Critic Model=VRPRM w/o RL2025.08 | 51.43 | — | — | — | — | |
| Qwen2.5-VL-32B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=GPT-OSS-120B-A5B2025.06 | 50.8 | — | — | — | — | |
| SwimBird2026.02 | 49.5 | — | — | — | — | |
| Claude-3.7-SonnetModel Category=Proprietary2025.06 | 49.3 | — | — | — | — | |
| Claude3.7-SonnetVLM Category=Proprietary VLMs2025.09 | 49.3 | — | — | — | — | |
| VL-Rethinker-72BModel Category=Open-Source2025.06 | 49.2 | — | — | — | — | |
| Qwen2.5-VL-72BModel Category=Open-Source2025.06 | 49.1 | — | — | — | — | |
| Qwen2.5-VL-72BReasoning Framework=Standard2025.06 | 49.1 | — | — | — | — | |
| Qwen2.5-VL-72BPolicy Model=Qwen2.5-VL-72B, Best-of-N Evaluation=82025.06 | 49.1 | — | — | — | — | |
| InternVL2.5-26B + VRPRM w/o RLBase Model=InternVL2.5-26B, Selection Strategy=Best-of-8, Critic Model=VRPRM w/o RL2025.08 | 48.76 | — | — | — | — | |
| InternVL3-38BModel Category=Open-Source2025.06 | 48.6 | — | — | — | — | |
| Qwen2.5-VL-32B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=Qwen3-8B2025.06 | 48.5 | — | — | — | — | |
| PyVision-ImageTooling Type=Dynamic Tooling2026.02 | 47.7 | — | — | — | — | |
| PyVision-RL-7BModel Category=Code-tool / visual-agent models2026.06 | 47.7 | — | — | — | — | |
| Gemini-2.0-FlashModel Category=Proprietary2025.06 | 47.4 | — | — | — | — | |
| Gemini-2.0-FlashSelection Strategy=Standard2025.08 | 47.4 | — | — | — | — | |
| Qwen2.5-VL-32BModel Category=Open-source MLLMs (no visual tool), Instruct variant=true2026.06 | 47.1 | — | — | — | — | |
| InternVL2.5-38B + VRPRMBase Model=InternVL2.5-38B, Selection Strategy=Best-of-8, Critic Model=VRPRM2025.08 | 46.86 | — | — | — | — | |
| Athena-PRMPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 46.4 | — | — | — | — | |
| InternVL2.5-38B + VisualPRMBase Model=InternVL2.5-38B, Selection Strategy=Best-of-8, Critic Model=VisualPRM2025.08 | 46.2 | — | — | — | — | |
| InternVL3-78BModel Category=Open-Source2025.06 | 46 | — | — | — | — | |
| GPT-4o-20241120Model Category=Proprietary2025.06 | 45.8 | — | — | — | — | |
| GPT-4oSelection Strategy=Standard2025.08 | 45.8 | — | — | — | — | |
| LUSPOBase model=Qwen2.5-VL-7B-Instruct, Algorithm=LUSPO2026.02 | 45.6 | — | — | — | — | |
| Qwen2.5-VL-7B w/ RAPIDReasoning Framework=RAPID, Reasoner Model=Qwen3-8B2025.06 | 45.4 | — | — | — | — | |
| Athena-ORMPolicy Model=Qwen2.5-VL-7B, Best-of-N Evaluation=82025.06 | 45.1 | — | — | — | — |