Visual Mathematical Reasoning on MathVerse VO
70.46BoN@8 AccuracyRL-Evol + Verifier (full VeriEvol)
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| RL-Evol + Verifier (full VeriEvol)Training Stage=RL, Model Size=7B, Evolution Strategy=Evolved, Verifier Usage=Yes (HTV-Agent)2026.06 | 70.46 | — | |
| RL-EvolTraining Stage=RL, Model Size=7B, Evolution Strategy=Evolved, Verifier Usage=No2026.06 | 69.67 | — | |
| RL-OriginTraining Stage=RL, Model Size=7B, Evolution Strategy=None, Verifier Usage=No2026.06 | 69.04 | — | |
| VeriEvol-SFT (VeriEvol-SFT-Init)Training Stage=SFT, Model Size=7B, Evolution Strategy=Evolved, Verifier Usage=No2026.06 | 67.01 | — | |
| Seed-only SFTTraining Stage=SFT, Model Size=7B, Evolution Strategy=None, Verifier Usage=No2026.06 | 65.02 | — | |
| OpenMMReasoner-7BTraining Stage=Baseline, Model Size=7B2026.06 | 63.8 | — | |
| MMR1-Math-v0Training Stage=Baseline, Model Size=7B2026.06 | 55.4 | — | |
| OVR-7BTraining Stage=Baseline, Model Size=7B2026.06 | 54.6 | — | |
| ReVisual-R1-7BTraining Stage=Baseline, Model Size=7B2026.06 | 53.6 | — | |
| VLAA-Thinker-7BTraining Stage=Baseline, Model Size=7B2026.06 | 48.2 | — | |
| Gemini-2.0-FlashReranking Strategy=Best-of-82026.03 | 47.8 | — | |
| InternVL2.5-38B + EVPV-PRMPolicy Model=InternVL2.5-38B, Process Reward Model (PRM)=EVPV-PRM, Reranking Strategy=Best-of-82026.03 | 47.67 | 10.77 | |
| InternVL2.5-38B + VisualPRMPolicy Model=InternVL2.5-38B, Process Reward Model (PRM)=VisualPRM, Reranking Strategy=Best-of-82026.03 | 46.7 | 9.8 | |
| VL-Rethinker-7BTraining Stage=Baseline, Model Size=7B2026.06 | 46.4 | — | |
| Claude-3.5-SonnetReranking Strategy=Best-of-82026.03 | 46.3 | — | |
| MM-Eureka-Qwen-7BTraining Stage=Baseline, Model Size=7B2026.06 | 45.4 | — | |
| WeThink-7BTraining Stage=Baseline, Model Size=7B2026.06 | 44.7 | — | |
| ThinkLite-VL-7BTraining Stage=Baseline, Model Size=7B2026.06 | 42.9 | — | |
| GPT-4oReranking Strategy=Best-of-82026.03 | 40.6 | — | |
| InternVL2.5-26B + VisualPRMPolicy Model=InternVL2.5-26B, Process Reward Model (PRM)=VisualPRM, Reranking Strategy=Best-of-82026.03 | 39.1 | 15.1 | |
| OpenVLThinker-7BTraining Stage=Baseline, Model Size=7B2026.06 | 38.1 | — | |
| InternVL2.5-38BPolicy Model=InternVL2.5-38B, Process Reward Model (PRM)=None, Reranking Strategy=Best-of-82026.03 | 36.9 | — | |
| InternVL2.5-8B + VisualPRMPolicy Model=InternVL2.5-8B, Process Reward Model (PRM)=VisualPRM, Reranking Strategy=Best-of-82026.03 | 35.8 | 13 | |
| Qwen2.5-VL-7B-InstructTraining Stage=Baseline, Model Size=7B2026.06 | 34.1 | — | |
| InternVL3-8BTraining Stage=Baseline, Model Size=8B2026.06 | 33.9 | — | |
| InternVL2.5-26B + EVPV-PRMPolicy Model=InternVL2.5-26B, Process Reward Model (PRM)=EVPV-PRM, Reranking Strategy=Best-of-82026.03 | 32.47 | 8.47 | |
| InternVL2.5-8B + EVPV-PRMPolicy Model=InternVL2.5-8B, Process Reward Model (PRM)=EVPV-PRM, Reranking Strategy=Best-of-82026.03 | 29.47 | 6.67 | |
| InternVL2.5-26BPolicy Model=InternVL2.5-26B, Process Reward Model (PRM)=None, Reranking Strategy=Best-of-82026.03 | 24 | — | |
| InternVL2.5-8BPolicy Model=InternVL2.5-8B, Process Reward Model (PRM)=None, Reranking Strategy=Best-of-82026.03 | 22.8 | — |