Hallucination Evaluation on AMBER Generative Subset (benign setting)
3.2CHAIR ScoreQwen-VL-Chat + ORCA
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Qwen-VL-Chat + ORCALVLM=Qwen-VL-Chat, Configuration=+ ORCA2025.09 | 3.2 | 8 | 0.4 | |
| Qwen-VL-Chat (Standalone)LVLM=Qwen-VL-Chat, Configuration=Standalone2025.09 | 7.5 | 40 | 4.1 | |
| MiniGPT-4 + ORCALVLM=MiniGPT-4, Configuration=+ ORCA2025.09 | 8.7 | 26 | 1.6 | |
| MiniGPT-4 (Standalone)LVLM=MiniGPT-4, Configuration=Standalone2025.09 | 16.3 | 84 | 12.6 | |
| mPlug-Owl + ORCALVLM=mPlug-Owl, Configuration=+ ORCA2025.09 | 19.7 | 56 | 7.7 | |
| mPlug-Owl (Standalone)LVLM=mPlug-Owl, Configuration=Standalone2025.09 | 23.2 | 82 | 16.7 |