Object Hallucination Mitigation on AMBER
12.1CHAIRICD
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| ICDBackbone Model=InstructBLIP2026.04 | 12.1 | 52.6 | 51.4 | 5.2 | |
| BaselineBackbone Model=InstructBLIP2026.04 | 11.6 | 53.4 | 51.7 | 5.3 | |
| BaselineBackbone Model=LLaVA1.52026.04 | 11.5 | 50.1 | 48.9 | 4.6 | |
| M3IDBackbone Model=InstructBLIP2026.04 | 11.5 | 52.5 | 51.4 | 4.6 | |
| VAF2026.04 | 10.9 | 33 | 28.1 | 3.7 | |
| Vanilla2026.04 | 10.6 | 50.9 | 36.4 | 4.1 | |
| VanillaBase Model=LLaVA-v1.52026.04 | 10.6 | 50.9 | 36.4 | 4.1 | |
| VCDBackbone Model=InstructBLIP2026.04 | 10.2 | 53.5 | 46.9 | 4.8 | |
| VCDBackbone Model=LLaVA1.52026.04 | 9.9 | 51.2 | 43.4 | 4.6 | |
| M3IDBackbone Model=LLaVA1.52026.04 | 9.8 | 55.6 | 48.4 | 3.6 | |
| ICD2026.04 | 9.3 | 50.5 | 41.3 | 4.9 | |
| ICDBackbone Model=LLaVA1.52026.04 | 9.1 | 51.2 | 40.6 | 4.3 | |
| FLBBackbone Model=InstructBLIP2026.04 | 9 | 53.6 | 43.8 | 4.7 | |
| VCD2026.04 | 8.9 | 51.8 | 41.9 | 4.3 | |
| MoDParadigm=Post-hoc hallucination mitigation, Base Model=LLaVA-1.5-7B2026.04 | 8.6 | 50.7 | 38.8 | 4.7 | |
| LLaVA-1.5-7B (Base)Paradigm=Standard LVLM2026.04 | 7.8 | 51 | 36.4 | 4.2 | |
| STICParadigm=Fine-tune LVLMs with self-improvement2026.04 | 7.6 | 52.1 | 35.8 | 4.4 | |
| Nullu2026.04 | 7.4 | 48.2 | 30.5 | 3.7 | |
| NulluBase Model=LLaVA-v1.52026.04 | 7.4 | 48.2 | 30.5 | 3.7 | |
| ICT2026.04 | 7 | 48 | 28.4 | 3.5 | |
| VTI2026.04 | 6.7 | 47.6 | 28.1 | 3.6 | |
| VTIBase Model=LLaVA-v1.52026.04 | 6.7 | 47.6 | 28.1 | 3.6 | |
| VCDParadigm=Post-hoc hallucination mitigation, Base Model=LLaVA-1.5-7B2026.04 | 6.7 | 46.5 | 27.8 | 2 | |
| MESA2026.04 | 6.4 | 50.4 | 27.4 | 3.2 | |
| MESABase Model=LLaVA-v1.52026.04 | 6.4 | 50.4 | 27.4 | 3.2 | |
| AVISCParadigm=Post-hoc hallucination mitigation, Base Model=LLaVA-1.5-7B2026.04 | 6.3 | 46.6 | 25.6 | 2 | |
| FLBBackbone Model=LLaVA1.52026.04 | 6.1 | 50.4 | 31.6 | 2.7 | |
| M3IDParadigm=Post-hoc hallucination mitigation, Base Model=LLaVA-1.5-7B2026.04 | 6 | 48.9 | 26 | 1.5 | |
| EOSParadigm=Fine-tune LVLMs with external annotation data2026.04 | 5.1 | 49.1 | 22.7 | 2 | |
| SENAParadigm=Fine-tune LVLMs with self-improvement2026.04 | 4.9 | 49.4 | 20.5 | 1.7 | |
| OctopusParadigm=Post-hoc hallucination mitigation, Base Model=LLaVA-1.5-7B2026.04 | 4.8 | 49.2 | 23.4 | 1.2 | |
| GPT-4VParadigm=Standard LVLM2026.04 | 4.6 | 67.1 | 30.7 | 2.6 | |
| MRGDParadigm=Post-hoc hallucination mitigation, Base Model=LLaVA-1.5-7B2026.04 | 4.4 | 62.5 | 21.9 | — | |
| PSRDParadigm=Post-hoc hallucination mitigation, Base Model=LLaVA-1.5-7B2026.04 | 3.9 | 48.2 | 20.1 | 2 | |
| CLIP-DPOParadigm=Fine-tune LVLMs with external annotation data2026.04 | 3.7 | 47.8 | 16.6 | 1.3 | |
| RLAIF-VParadigm=Fine-tune LVLMs with external annotation data2026.04 | 2.9 | 50.2 | 16 | 1 | |
| LLaVA-DPOParadigm=Fine-tune LVLMs with external annotation data2026.04 | 2.8 | 47.8 | 15.5 | 1.6 | |
| HSA-DPOParadigm=Fine-tune LVLMs with external annotation data2026.04 | 2.1 | 47.3 | 13.4 | 1.2 |