Vision-Language Hallucination on MMHal-Bench
32HalQwen2.5-VL-7B + P2-DPO
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Qwen2.5-VL-7B + P2-DPOModel=Qwen2.5-VL-7B, Pairs Source=Self2026.06 | 32 | 3.94 | |
| Qwen2.5-VL-7B + DPORLHF-VModel=Qwen2.5-VL-7B, Pairs Source=Human2026.06 | 34.5 | 3.63 | |
| Qwen2.5-VL-7BModel=Qwen2.5-VL-7B, Pairs Source=Base2026.06 | 34.9 | 3.47 | |
| Qwen2.5-VL-3B + P2-DPOModel=Qwen2.5-VL-3B, Pairs Source=Self2026.06 | 39 | 3.49 | |
| Qwen2.5-VL-3B + DPORLHF-VModel=Qwen2.5-VL-3B, Pairs Source=Human2026.06 | 40 | 3.41 | |
| Qwen2.5-VL-3BModel=Qwen2.5-VL-3B, Pairs Source=Base2026.06 | 42 | 3.38 | |
| LLaVA-1.5-7B + V-DPOSADModel=LLaVA-1.5-7B, Pairs Source=AI2026.06 | 53 | 2.36 | |
| LLaVA-1.5-7B + V-DPORLHF-VModel=LLaVA-1.5-7B, Pairs Source=Human2026.06 | 56 | 2.16 | |
| LLaVA-1.5-7B + P2-DPOModel=LLaVA-1.5-7B, Pairs Source=Self2026.06 | 56 | 2.43 | |
| LLaVA-1.5-7B + VCDModel=LLaVA-1.5-7B, Pairs Source=None2026.06 | 58 | 2.12 | |
| LLaVA-1.5-7B + HA-DPOModel=LLaVA-1.5-7B, Pairs Source=AI2026.06 | 60 | 1.97 | |
| LLaVA-1.5-7B + DPORLHF-VModel=LLaVA-1.5-7B, Pairs Source=Human2026.06 | 60 | 2.08 | |
| LLaVA-1.5-7BModel=LLaVA-1.5-7B, Pairs Source=Base2026.06 | 62 | 1.97 |