Expert-level multidisciplinary QA on MMMU (dev val)
52.8Pass@1Vision-SR1
Evaluation Results
| Method | Links | |
|---|---|---|
| Vision-SR1Model Category=Reasoning-Oriented2025.12 | 52.8 | |
| R1-ShareVL-7BModel Category=Reasoning-Oriented2025.12 | 52.28 | |
| PeRL-VLBase Model=Qwen2.5-VL-7B-Instruct2025.12 | 52.22 | |
| RL (Conditional Hard Gate)Training Stage=PeRL: Perception Stage Only, Reward Function (γ)=02025.12 | 52.11 | |
| RL (Conditional Soft Gate)Training Stage=PeRL: Perception Stage Only, Reward Function (γ)=0.52025.12 | 52.09 | |
| VL-CogitoModel Category=Reasoning-Oriented2025.12 | 51.72 | |
| PeBR-R1-7BModel Category=Reasoning-Oriented2025.12 | 51.59 | |
| RL (Verifiable Rewards)Training Stage=PeRL: Perception Stage Only, Reward Function (γ)=12025.12 | 51.42 | |
| ThinkLite-VLModel Category=Reasoning-Oriented2025.12 | 50.81 | |
| RL (Aggregated Rewards)Training Stage=PeRL: Perception Stage Only2025.12 | 50.8 | |
| SFT (OpenThought)Training Stage=PeRL: Reasoning Stage Only2025.12 | 48.51 | |
| Qwen2.5-VL-7BModel Category=Base2025.12 | 48.15 | |
| SFT (GPT-4o)Training Stage=PeRL: Reasoning Stage Only2025.12 | 48.06 |