Multimodal Reasoning on MathVerse (mean@8 acc %)
60.95Mean@8 AccuracyDAPO
Evaluation Results
| Method | Links | |
|---|---|---|
| DAPOBase Model=Qwen2-VL-8B Instruct, Temperature=0.62026.05 | 60.95 | |
| PAPO_DBase Model=Qwen2-VL-8B Instruct, Temperature=0.62026.05 | 59.64 | |
| VPPO_DBase Model=Qwen2-VL-8B Instruct, Temperature=0.62026.05 | 59.64 | |
| RAPO_DBase Model=Qwen2-VL-8B Instruct, Temperature=0.62026.05 | 58.5 | |
| RAPOBase Model=Qwen3-VL-8B-Instruct2026.05 | 56.72 | |
| VPPOBase Model=Qwen3-VL-8B-Instruct2026.05 | 56.35 | |
| GRPOBase Model=Qwen3-VL-8B-Instruct2026.05 | 55.23 | |
| PAPOBase Model=Qwen3-VL-8B-Instruct2026.05 | 54.57 | |
| BaseBase Model=Qwen3-VL-8B-Instruct2026.05 | 50.21 | |
| Base (Qwen2-VL-8B Instruct)Temperature=0.62026.05 | 50.21 | |
| RAPO_DBase Model=Qwen2.5-VL-7B-Instruct2026.05 | 47.59 | |
| RAPO_GBase Model=Qwen2.5-VL-7B-Instruct2026.05 | 47.34 | |
| R1-ShareVLBase Model=Qwen2.5-VL-7B-Instruct2026.05 | 46.22 | |
| RAPOBase Model=Qwen3-VL-2B-Instruct2026.05 | 46.13 | |
| VPPOBase Model=Qwen3-VL-2B-Instruct2026.05 | 45.43 | |
| VL-RethinkerBase Model=Qwen2.5-VL-7B-Instruct2026.05 | 45.3 | |
| PAPOBase Model=Qwen3-VL-2B-Instruct2026.05 | 45.18 | |
| MM-EurekaBase Model=Qwen2.5-VL-7B-Instruct2026.05 | 45.18 | |
| GRPOBase Model=Qwen3-VL-2B-Instruct2026.05 | 44.35 | |
| BaseBase Model=Qwen3-VL-2B-Instruct2026.05 | 36.75 | |
| RAPO_DBase Model=Qwen2-VL-7B-Instruct2026.05 | 36.31 | |
| RAPO_GBase Model=Qwen2-VL-7B-Instruct2026.05 | 32.99 | |
| BaseBase Model=Qwen2.5-VL-7B-Instruct2026.05 | 29.44 | |
| MINT-CoTBase Model=Qwen2-VL-7B-Instruct2026.05 | 24.62 | |
| TVCBase Model=Qwen2-VL-7B-Instruct2026.05 | 21.98 | |
| BaseBase Model=Qwen2-VL-7B-Instruct2026.05 | 10.25 |