Object Hallucination Evaluation on POPE Concurrence
92.33AccuracyORCA
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| ORCABase LVLM=Qwen-VL-Chat (13B), Agent Backbone=LLaMA-3.2-Vision-Instruct (11B), Visual Backbone=OpenCLIP2025.09 | 92.33 | 91.93 | |
| ORCABase LVLM=MiniGPT-4 (13B), Agent Backbone=LLaMA-3.2-Vision-Instruct (11B), Visual Backbone=EVA2025.09 | 92 | 91.84 | |
| ORCABase LVLM=mPlug-Owl (7B), Agent Backbone=LLaMA-3.2-Vision-Instruct (11B), Visual Backbone=CLIP2025.09 | 90.33 | 90.24 | |
| mPlug-Owl + ORCALVLM=mPlug-Owl, Variant=+ ORCA2025.09 | 86 | 85.21 | |
| Qwen-VL-ChatParameters=13B, Backbone=OpenCLIP, Setup=Standalone2025.09 | 85.67 | 85.71 | |
| MiniGPT-4 + ORCALVLM=MiniGPT-4, Variant=+ ORCA2025.09 | 84.67 | 83.57 | |
| Qwen-VL-Chat + ORCALVLM=Qwen-VL-Chat, Variant=+ ORCA2025.09 | 83.67 | 80.93 | |
| Qwen-VL-Chat (Standalone)LVLM=Qwen-VL-Chat, Variant=Standalone2025.09 | 82.33 | 79.21 | |
| mPlug-Owl + ORCALVLM=mPlug-Owl, Method=+ ORCA, Attack Scenario=Adversarial Illusions2025.09 | 81.33 | 80 | |
| MiniGPT-4 + ORCALVLM=MiniGPT-4, Method=+ ORCA, Attack Scenario=Adversarial Illusions2025.09 | 80.67 | 79.58 | |
| Qwen-VL-Chat (Standalone)LVLM=Qwen-VL-Chat, Method=Standalone, Attack Scenario=Adversarial Illusions2025.09 | 79.67 | 76.62 | |
| Qwen-VL-Chat + ORCALVLM=Qwen-VL-Chat, Method=+ ORCA, Attack Scenario=Adversarial Illusions2025.09 | 77 | 72.06 | |
| MiniGPT-4Parameters=13B, Backbone=EVA, Setup=Standalone2025.09 | 70 | 74.29 | |
| MiniGPT-4 (Standalone)LVLM=MiniGPT-4, Variant=Standalone2025.09 | 64.33 | 73.58 | |
| MiniGPT-4 (Standalone)LVLM=MiniGPT-4, Method=Standalone, Attack Scenario=Adversarial Illusions2025.09 | 61.67 | 71.18 | |
| mPlug-OwlParameters=7B, Backbone=CLIP, Setup=Standalone2025.09 | 51 | 67.11 | |
| mPlug-Owl (Standalone)LVLM=mPlug-Owl, Variant=Standalone2025.09 | 50 | 66.67 | |
| mPlug-Owl (Standalone)LVLM=mPlug-Owl, Method=Standalone, Attack Scenario=Adversarial Illusions2025.09 | 50 | 66.67 |