Visual Mathematical Reasoning on DynaMath (Worst-case Accuracy)
35.73Worst-case AccuracyRL-Evol + Verifier (full VeriEvol)
Evaluation Results
| Method | Links | |
|---|---|---|
| RL-Evol + Verifier (full VeriEvol)Training Stage=RL, Model Size=7B, Evolution Strategy=Evolved, Verifier Usage=Yes (HTV-Agent)2026.06 | 35.73 | |
| OpenMMReasoner-7BTraining Stage=Baseline, Model Size=7B2026.06 | 34.9 | |
| OVR-7BTraining Stage=Baseline, Model Size=7B2026.06 | 33.5 | |
| RL-EvolTraining Stage=RL, Model Size=7B, Evolution Strategy=Evolved, Verifier Usage=No2026.06 | 32.14 | |
| VeriEvol-SFT (VeriEvol-SFT-Init)Training Stage=SFT, Model Size=7B, Evolution Strategy=Evolved, Verifier Usage=No2026.06 | 30.94 | |
| RL-OriginTraining Stage=RL, Model Size=7B, Evolution Strategy=None, Verifier Usage=No2026.06 | 30.54 | |
| MMR1-Math-v0Training Stage=Baseline, Model Size=7B2026.06 | 27.9 | |
| ReVisual-R1-7BTraining Stage=Baseline, Model Size=7B2026.06 | 27.5 | |
| Seed-only SFTTraining Stage=SFT, Model Size=7B, Evolution Strategy=None, Verifier Usage=No2026.06 | 26.96 | |
| WeThink-7BTraining Stage=Baseline, Model Size=7B2026.06 | 24.4 | |
| MM-Eureka-Qwen-7BTraining Stage=Baseline, Model Size=7B2026.06 | 23 | |
| InternVL3-8BTraining Stage=Baseline, Model Size=8B2026.06 | 23 | |
| VLAA-Thinker-7BTraining Stage=Baseline, Model Size=7B2026.06 | 22.4 | |
| Qwen2.5-VL-7B-InstructTraining Stage=Baseline, Model Size=7B2026.06 | 18 | |
| VL-Rethinker-7BTraining Stage=Baseline, Model Size=7B2026.06 | 17.8 | |
| OpenVLThinker-7BTraining Stage=Baseline, Model Size=7B2026.06 | 16.8 | |
| ThinkLite-VL-7BTraining Stage=Baseline, Model Size=7B2026.06 | 16.5 |