Multi-modal Understanding on MMMU Pro (Overall)
53.3ScorePRISM + GRPO
Evaluation Results
| Method | Links | |
|---|---|---|
| PRISM + GRPOBase Model=Qwen3-VL-8B2026.04 | 53.3 | |
| PRISM + GSPOBase Model=Qwen3-VL-8B2026.04 | 52.7 | |
| PRISM + DAPOBase Model=Qwen3-VL-8B2026.04 | 52.4 | |
| InstructBase Model=Qwen3-VL-8B2026.04 | 52.3 | |
| PRISM + GSPOBase Model=Qwen3-VL-4B2026.04 | 51.1 | |
| PRISM + DAPOBase Model=Qwen3-VL-4B2026.04 | 50.4 | |
| PRISM + GRPOBase Model=Qwen3-VL-4B2026.04 | 49.7 | |
| + DAPOBase Model=Qwen3-VL-8B2026.04 | 49 | |
| + GRPOBase Model=Qwen3-VL-8B2026.04 | 48.8 | |
| + DAPOBase Model=Qwen3-VL-4B2026.04 | 48 | |
| + GSPOBase Model=Qwen3-VL-8B2026.04 | 47.8 | |
| + GRPOBase Model=Qwen3-VL-4B2026.04 | 47.3 | |
| + GSPOBase Model=Qwen3-VL-4B2026.04 | 45.6 | |
| InstructBase Model=Qwen3-VL-4B2026.04 | 45.1 | |
| PRISMBase Model=Qwen3-VL-8B2026.04 | 43.4 | |
| + SFTBase Model=Qwen3-VL-8B2026.04 | 42.9 | |
| + SFTBase Model=Qwen3-VL-4B2026.04 | 42.8 | |
| PRISMBase Model=Qwen3-VL-4B2026.04 | 42.8 | |
| VanillaMode=Upper bound2024.12 | 38.3 | |
| SparseVLMImage token reduction ratio=66.7%2024.12 | 38.3 | |
| VisionZipImage token reduction ratio=66.7%2024.12 | 38.2 | |
| iLLaVAImage token reduction ratio=66.7%2024.12 | 38.1 | |
| PyramidDropImage token reduction ratio=66.7%2024.12 | 37.8 | |
| iLLaVAImage token reduction ratio=77.8%2024.12 | 37.8 | |
| FasterVLMImage token reduction ratio=66.7%2024.12 | 37.7 | |
| VisionZipImage token reduction ratio=77.8%2024.12 | 37.5 | |
| SparseVLMImage token reduction ratio=77.8%2024.12 | 36.6 | |
| iLLaVAImage token reduction ratio=88.9%2024.12 | 36.6 | |
| FasterVLMImage token reduction ratio=77.8%2024.12 | 36.5 | |
| PyramidDropImage token reduction ratio=77.8%2024.12 | 36.4 | |
| FasterVLMImage token reduction ratio=88.9%2024.12 | 36.4 | |
| SparseVLMImage token reduction ratio=88.9%2024.12 | 36.2 | |
| VisionZipImage token reduction ratio=88.9%2024.12 | 35.9 | |
| PyramidDropImage token reduction ratio=88.9%2024.12 | 34 |