General Multimodal Understanding on MMMU Pro
56.9AccuracyQwen3-VL-32B-Instruct
Evaluation Results
| Method | Links | |
|---|---|---|
| Qwen3-VL-32B-InstructBackbone model=Qwen3-VL-32B-Instruct2026.05 | 56.9 | |
| IVR-R1Backbone model=Qwen3-VL-4B2026.05 | 53.3 | |
| Vision-R1Backbone model=Qwen3-VL-4B2026.05 | 52.5 | |
| Vision-SR1Backbone model=Qwen3-VL-4B2026.05 | 49.3 | |
| Supervised Fine-tuning (before RL)Backbone model=Qwen3-VL-4B2026.05 | 48.6 | |
| IVR-R1Backbone model=Qwen2.5-VL-7B2026.05 | 47.8 | |
| Vision-R1Backbone model=Qwen2.5-VL-7B2026.05 | 47.2 | |
| Vision-SR1Backbone model=Qwen2.5-VL-7B2026.05 | 45.5 | |
| Supervised Fine-tuning (before RL)Backbone model=Qwen2.5-VL-7B2026.05 | 43.6 | |
| Zero-shot Inference (before RL)Backbone model=Qwen3-VL-4B2026.05 | 42.7 | |
| VPPOModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 38.8 | |
| PRPOModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 38.8 | |
| VL-Rethinker-7BModel Scale=7B, Rollout=82026.06 | 37.1 | |
| Qwen2.5-VL-72B-InstructBackbone model=Qwen2.5-VL-72B-Instruct2026.05 | 36.6 | |
| PAPO-DModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 36.3 | |
| PAPO-GModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 35.5 | |
| GRPOModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 35.2 | |
| R1-ShareVL-7BModel Scale=7B, Rollout=82026.06 | 35.1 | |
| NoisyRollout-7BModel Scale=7B, Rollout=82026.06 | 34.5 | |
| PRPOModel Scale=3B, Base Model=Qwen2.5-VL-3B, Rollout=82026.06 | 30.3 | |
| MM-Eureka-7BModel Scale=7B, Rollout=82026.06 | 30.3 | |
| DAPOModel Scale=7B, Base Model=Qwen2.5-VL-7B, Rollout=82026.06 | 29 | |
| PAPO-DModel Scale=3B, Base Model=Qwen2.5-VL-3B, Rollout=82026.06 | 28.8 | |
| VPPOModel Scale=3B, Base Model=Qwen2.5-VL-3B, Rollout=82026.06 | 28.8 | |
| DAPOModel Scale=3B, Base Model=Qwen2.5-VL-3B, Rollout=82026.06 | 28.4 | |
| PAPO-GModel Scale=3B, Base Model=Qwen2.5-VL-3B, Rollout=82026.06 | 26.4 | |
| GRPOModel Scale=3B, Base Model=Qwen2.5-VL-3B, Rollout=82026.06 | 25.7 | |
| Qwen2.5-VL-7BModel Scale=7B, Rollout=82026.06 | 25.2 | |
| Qwen2.5-VL-32B-InstructBackbone model=Qwen2.5-VL-32B-Instruct2026.05 | 21.7 | |
| Qwen2.5-VL-3BModel Scale=3B, Rollout=82026.06 | 19.3 | |
| Zero-shot Inference (before RL)Backbone model=Qwen2.5-VL-7B2026.05 | 16.2 |