Visual Mathematical Reasoning on MathVista MINI (Pass@1)
70.5Pass@1PeBR-R1-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| PeBR-R1-7BModel Category=Reasoning-Oriented2025.12 | 70.5 | |
| SFT (OpenThought)Training Stage=PeRL: Reasoning Stage Only2025.12 | 67.5 | |
| PeRL-VLBase Model=Qwen2.5-VL-7B-Instruct2025.12 | 67.05 | |
| RL (Conditional Hard Gate)Training Stage=PeRL: Perception Stage Only, Reward Function (γ)=02025.12 | 66.11 | |
| R1-ShareVL-7BModel Category=Reasoning-Oriented2025.12 | 66.1 | |
| RL (Conditional Soft Gate)Training Stage=PeRL: Perception Stage Only, Reward Function (γ)=0.52025.12 | 65.8 | |
| RL (Aggregated Rewards)Training Stage=PeRL: Perception Stage Only2025.12 | 65.45 | |
| ThinkLite-VLModel Category=Reasoning-Oriented2025.12 | 65 | |
| RL (Verifiable Rewards)Training Stage=PeRL: Perception Stage Only, Reward Function (γ)=12025.12 | 64.15 | |
| Vision-SR1Model Category=Reasoning-Oriented2025.12 | 63.5 | |
| VL-CogitoModel Category=Reasoning-Oriented2025.12 | 61.3 | |
| SFT (GPT-4o)Training Stage=PeRL: Reasoning Stage Only2025.12 | 58.9 | |
| Qwen2.5-VL-7BModel Category=Base2025.12 | 57.78 |