Mathematical & Geometric Reasoning on DynaMath (accuracy@8)
73.1Accuracy@8Qwen2.5-VL-32B + VPPO
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-VL-32B + VPPOModel Scale=32B, Algorithm=VPPO2025.10 | 73.1 | |
| Qwen2.5-VL-32B + DAPOModel Scale=32B, Algorithm=DAPO2025.10 | 72.6 | |
| NoisyRollout-32BModel Scale=32B, Prompting Source=Official author-provided, Note=Trained using the training set of Geo3k2025.10 | 72.2 | |
| MM-Eureka-32BModel Scale=32B, Prompting Source=Official author-provided2025.10 | 72 | |
| Qwen2.5-VL-32B + GRPOModel Scale=32B, Algorithm=GRPO2025.10 | 71.6 | |
| Qwen2.5-VL-32BModel Scale=32B2025.10 | 68.7 | |
| Qwen2.5-VL-7B + VPPOModel Scale=7B, Algorithm=VPPO2025.10 | 68.1 | |
| PAPO-D-7BModel Scale=7B, Training Strategy=Pure RL2025.10 | 66.8 | |
| Qwen2.5-VL-7B + DAPOModel Scale=7B, Algorithm=DAPO2025.10 | 66.6 | |
| Qwen2.5-VL-7B + GRPOModel Scale=7B, Algorithm=GRPO2025.10 | 65.8 | |
| VL-Rethinker-7BModel Scale=7B, Training Strategy=Pure RL, Prompting Source=Official author-provided2025.10 | 65.7 | |
| NoisyRollout-7BModel Scale=7B, Training Strategy=Pure RL, Prompting Source=Official author-provided, Note=Trained using the training set of Geo3k2025.10 | 65.5 | |
| MM-Eureka-7BModel Scale=7B, Training Strategy=Pure RL, Prompting Source=Official author-provided2025.10 | 65.4 | |
| R1-ShareVL-7BModel Scale=7B, Training Strategy=Pure RL, Prompting Source=Official author-provided2025.10 | 65.1 | |
| ThinkLite-7BModel Scale=7B, Training Strategy=Pure RL, Prompting Source=Official author-provided2025.10 | 64.6 | |
| Qwen2.5-VL-7BModel Scale=7B2025.10 | 55.7 |