Multimodal Reasoning on MathVision (mean@8 acc %)
51.36Mean@8 AccuracyRAPO_D
Evaluation Results
| Method | Links | |
|---|---|---|
| RAPO_DBase Model=Qwen2-VL-8B Instruct, Temperature=0.62026.05 | 51.36 | |
| VPPO_DBase Model=Qwen2-VL-8B Instruct, Temperature=0.62026.05 | 50.99 | |
| DAPOBase Model=Qwen2-VL-8B Instruct, Temperature=0.62026.05 | 50.33 | |
| PAPO_DBase Model=Qwen2-VL-8B Instruct, Temperature=0.62026.05 | 49.67 | |
| RAPOBase Model=Qwen3-VL-8B-Instruct2026.05 | 48.27 | |
| GRPOBase Model=Qwen3-VL-8B-Instruct2026.05 | 48.03 | |
| VPPOBase Model=Qwen3-VL-8B-Instruct2026.05 | 48.03 | |
| PAPOBase Model=Qwen3-VL-8B-Instruct2026.05 | 43.75 | |
| BaseBase Model=Qwen3-VL-8B-Instruct2026.05 | 41.74 | |
| Base (Qwen2-VL-8B Instruct)Temperature=0.62026.05 | 41.74 | |
| VL-RethinkerBase Model=Qwen2.5-VL-7B-Instruct2026.05 | 34.21 | |
| RAPOBase Model=Qwen3-VL-2B-Instruct2026.05 | 33.55 | |
| MM-EurekaBase Model=Qwen2.5-VL-7B-Instruct2026.05 | 32.57 | |
| VPPOBase Model=Qwen3-VL-2B-Instruct2026.05 | 30.59 | |
| RAPO_DBase Model=Qwen2.5-VL-7B-Instruct2026.05 | 30.26 | |
| GRPOBase Model=Qwen3-VL-2B-Instruct2026.05 | 28.29 | |
| PAPOBase Model=Qwen3-VL-2B-Instruct2026.05 | 28.29 | |
| RAPO_GBase Model=Qwen2.5-VL-7B-Instruct2026.05 | 27.96 | |
| BaseBase Model=Qwen3-VL-2B-Instruct2026.05 | 27.3 | |
| R1-ShareVLBase Model=Qwen2.5-VL-7B-Instruct2026.05 | 25.99 | |
| RAPO_DBase Model=Qwen2-VL-7B-Instruct2026.05 | 23.36 | |
| RAPO_GBase Model=Qwen2-VL-7B-Instruct2026.05 | 22.37 | |
| MINT-CoTBase Model=Qwen2-VL-7B-Instruct2026.05 | 22.04 | |
| BaseBase Model=Qwen2.5-VL-7B-Instruct2026.05 | 21.05 | |
| TVCBase Model=Qwen2-VL-7B-Instruct2026.05 | 18.75 | |
| BaseBase Model=Qwen2-VL-7B-Instruct2026.05 | 15.79 |