Visual Mathematical Reasoning on MMK12 (val)
58.86AccuracyQwen2.5-VL-3B-Instruct + VEPO
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen2.5-VL-3B-Instruct + VEPOBase Model Size=3B, Fine-tuning=VEPO2026.06 | 58.86 | |
| PAPO-DAPO-3BBase Model Size=3B2026.06 | 58.85 | |
| VPPO-3BBase Model Size=3B2026.06 | 58.24 | |
| NoisyRollout-3BBase Model Size=3B2026.06 | 58.12 | |
| Qwen2.5-VL-3B-Instruct + GRPOBase Model Size=3B, Fine-tuning=GRPO2026.06 | 57.37 | |
| Qwen2.5-VL-3B-Instruct + GRPO (Top 20% high entropy)Base Model Size=3B, Fine-tuning=GRPO, Constraint=Top 20% high entropy2026.06 | 56.81 | |
| R1-ShareVLBase Model Size=3B2026.06 | 56.26 | |
| Qwen2.5-VL-3B-Instruct + GRPO (Top 40% high entropy)Base Model Size=3B, Fine-tuning=GRPO, Constraint=Top 40% high entropy2026.06 | 54.52 | |
| Qwen2.5-VL-3B-InstructBase Model Size=3B2026.06 | 45.35 |