Multimodal Reasoning on MMMU-Pro (mean@8 acc %)
51.1MMMU-Pro mean@8 AccRAPO
Evaluation Results
| Method | Links | |
|---|---|---|
| RAPOBase Model=Qwen3-VL-8B-Instruct2026.05 | 51.1 | |
| PAPOBase Model=Qwen3-VL-8B-Instruct2026.05 | 50 | |
| VPPOBase Model=Qwen3-VL-8B-Instruct2026.05 | 49.16 | |
| GRPOBase Model=Qwen3-VL-8B-Instruct2026.05 | 48.5 | |
| BaseBase Model=Qwen3-VL-8B-Instruct2026.05 | 37.59 | |
| RAPO_DBase Model=Qwen2.5-VL-7B-Instruct2026.05 | 36.04 | |
| RAPO_GBase Model=Qwen2.5-VL-7B-Instruct2026.05 | 36.01 | |
| R1-ShareVLBase Model=Qwen2.5-VL-7B-Instruct2026.05 | 34.45 | |
| VL-RethinkerBase Model=Qwen2.5-VL-7B-Instruct2026.05 | 34.1 | |
| RAPOBase Model=Qwen3-VL-2B-Instruct2026.05 | 32.2 | |
| VPPOBase Model=Qwen3-VL-2B-Instruct2026.05 | 31.73 | |
| PAPOBase Model=Qwen3-VL-2B-Instruct2026.05 | 31.45 | |
| GRPOBase Model=Qwen3-VL-2B-Instruct2026.05 | 30.61 | |
| RAPO_DBase Model=Qwen2-VL-7B-Instruct2026.05 | 29.39 | |
| RAPO_GBase Model=Qwen2-VL-7B-Instruct2026.05 | 28.22 | |
| MM-EurekaBase Model=Qwen2.5-VL-7B-Instruct2026.05 | 26.82 | |
| BaseBase Model=Qwen3-VL-2B-Instruct2026.05 | 24.44 | |
| TVCBase Model=Qwen2-VL-7B-Instruct2026.05 | 23.34 | |
| BaseBase Model=Qwen2.5-VL-7B-Instruct2026.05 | 21.22 | |
| MINT-CoTBase Model=Qwen2-VL-7B-Instruct2026.05 | 17.63 | |
| BaseBase Model=Qwen2-VL-7B-Instruct2026.05 | 11.73 |