General-purpose multiple-choice evaluation on MMBench EN V11 (dev)
84.3Pass@1Vision-SR1
Evaluation Results
| Method | Links | |
|---|---|---|
| Vision-SR1Model Category=Reasoning-Oriented2025.12 | 84.3 | |
| RL (Conditional Hard Gate)Training Stage=PeRL: Perception Stage Only, Reward Function (γ)=02025.12 | 84.29 | |
| RL (Conditional Soft Gate)Training Stage=PeRL: Perception Stage Only, Reward Function (γ)=0.52025.12 | 84.15 | |
| VL-CogitoModel Category=Reasoning-Oriented2025.12 | 83.92 | |
| PeBR-R1-7BModel Category=Reasoning-Oriented2025.12 | 83.78 | |
| R1-ShareVL-7BModel Category=Reasoning-Oriented2025.12 | 83.67 | |
| PeRL-VLBase Model=Qwen2.5-VL-7B-Instruct2025.12 | 83.5 | |
| RL (Verifiable Rewards)Training Stage=PeRL: Perception Stage Only, Reward Function (γ)=12025.12 | 83.2 | |
| RL (Aggregated Rewards)Training Stage=PeRL: Perception Stage Only2025.12 | 83.2 | |
| Qwen2.5-VL-7BModel Category=Base2025.12 | 82.9 | |
| ThinkLite-VLModel Category=Reasoning-Oriented2025.12 | 82.65 | |
| SFT (GPT-4o)Training Stage=PeRL: Reasoning Stage Only2025.12 | 82.57 | |
| SFT (OpenThought)Training Stage=PeRL: Reasoning Stage Only2025.12 | 81.1 |