Hallucination on HallusionBench (Pass@K Metrics)
74Pass@1GRPO
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GRPOevaluation_mode=Out-of-Distribution (OOD), training_protocol=Transfer from ChartQA2026.02 | 74 | 94 | — | |
| MIGevaluation_mode=Out-of-Distribution (OOD), training_protocol=Transfer from ChartQA2026.02 | 71 | 97 | -3 | |
| Base Modelevaluation_mode=Out-of-Distribution (OOD), training_protocol=Transfer from ChartQA2026.02 | 69 | 98 | — | |
| PeRL-VLBase Model=Qwen2.5-VL-7B-Instruct2025.12 | 55.9 | — | — | |
| SFT (GPT-4o)Training Stage=PeRL: Reasoning Stage Only2025.12 | 52.3 | — | — | |
| PeBR-R1-7BModel Category=Reasoning-Oriented2025.12 | 49.56 | — | — | |
| RL (Aggregated Rewards)Training Stage=PeRL: Perception Stage Only2025.12 | 47.33 | — | — | |
| RL (Conditional Hard Gate)Training Stage=PeRL: Perception Stage Only, Reward Function (γ)=02025.12 | 47.22 | — | — | |
| VL-CogitoModel Category=Reasoning-Oriented2025.12 | 46.63 | — | — | |
| RL (Conditional Soft Gate)Training Stage=PeRL: Perception Stage Only, Reward Function (γ)=0.52025.12 | 46.12 | — | — | |
| R1-ShareVL-7BModel Category=Reasoning-Oriented2025.12 | 44.53 | — | — | |
| Qwen2.5-VL-7BModel Category=Base2025.12 | 43.8 | — | — | |
| Vision-SR1Model Category=Reasoning-Oriented2025.12 | 43.2 | — | — | |
| SFT (OpenThought)Training Stage=PeRL: Reasoning Stage Only2025.12 | 41.4 | — | — | |
| RL (Verifiable Rewards)Training Stage=PeRL: Perception Stage Only, Reward Function (γ)=12025.12 | 37.67 | — | — | |
| ThinkLite-VLModel Category=Reasoning-Oriented2025.12 | 36.1 | — | — |