OCR-centric visual reasoning on OCRBench
86.56Pass@1RL (Conditional Soft Gate)
Evaluation Results
| Method | Links | |
|---|---|---|
| RL (Conditional Soft Gate)Training Stage=PeRL: Perception Stage Only, Reward Function (γ)=0.52025.12 | 86.56 | |
| PeBR-R1-7BModel Category=Reasoning-Oriented2025.12 | 86.31 | |
| RL (Aggregated Rewards)Training Stage=PeRL: Perception Stage Only2025.12 | 86.3 | |
| RL (Conditional Hard Gate)Training Stage=PeRL: Perception Stage Only, Reward Function (γ)=02025.12 | 86.16 | |
| R1-ShareVL-7BModel Category=Reasoning-Oriented2025.12 | 85.85 | |
| RL (Verifiable Rewards)Training Stage=PeRL: Perception Stage Only, Reward Function (γ)=12025.12 | 85.13 | |
| PeRL-VLBase Model=Qwen2.5-VL-7B-Instruct2025.12 | 85.1 | |
| ThinkLite-VLModel Category=Reasoning-Oriented2025.12 | 84.82 | |
| VL-CogitoModel Category=Reasoning-Oriented2025.12 | 84.64 | |
| SFT (GPT-4o)Training Stage=PeRL: Reasoning Stage Only2025.12 | 82.58 | |
| Vision-SR1Model Category=Reasoning-Oriented2025.12 | 82.37 | |
| SFT (OpenThought)Training Stage=PeRL: Reasoning Stage Only2025.12 | 82.2 | |
| Qwen2.5-VL-7BModel Category=Base2025.12 | 80.2 |