Hallucination Assessment on AMBER
10.6CHAIR_sSampling
Evaluation Results
| Method | Links | |||||||
|---|---|---|---|---|---|---|---|---|
| SamplingBackbone=LLaVA-1.5 (7B), Prompt=Please describe this image in detail2026.01 | 10.6 | — | 44.3 | 4 | — | — | — | |
| SamplingBackbone=LLaVA-NeXT, Prompt=Please describe this image in detail2026.01 | 10.5 | — | 57.1 | 4.1 | — | — | — | |
| ClearSightBackbone=LLaVA-1.5 (7B), Prompt=Please describe this image in detail2026.01 | 10.5 | — | 44.1 | 4.2 | — | — | — | |
| ICDBackbone=LLaVA-NeXT, Prompt=Please describe this image in detail2026.01 | 10.1 | — | 53 | 4.5 | — | — | — | |
| ICDBackbone=LLaVA-1.5 (7B), Prompt=Please describe this image in detail2026.01 | 10 | — | 44.8 | 4.3 | — | — | — | |
| OPERABackbone=LLaVA-1.5 (7B), Prompt=Please describe this image in detail2026.01 | 9.8 | — | 43 | 4.5 | — | — | — | |
| OPERABackbone=LLaVA-NeXT, Prompt=Please describe this image in detail2026.01 | 9.8 | — | 51 | 4.3 | — | — | — | |
| VCDBackbone=LLaVA-NeXT, Prompt=Please describe this image in detail2026.01 | 9.5 | — | 55.7 | 4.2 | — | — | — | |
| SamplingBackbone=LLaVA-1.5 (13B), Prompt=Please describe this image in detail2026.01 | 9.3 | — | 41.3 | 4.2 | — | — | — | |
| SIDBackbone=LLaVA-1.5 (7B), Prompt=Please describe this image in detail2026.01 | 9.3 | — | 43.7 | 3.7 | — | — | — | |
| ClearSightBackbone=LLaVA-1.5 (13B), Prompt=Please describe this image in detail2026.01 | 9.2 | — | 40.5 | 4 | — | — | — | |
| ClearSightBackbone=LLaVA-NeXT, Prompt=Please describe this image in detail2026.01 | 9.1 | — | 54.4 | 4.5 | — | — | — | |
| SIDBackbone=LLaVA-NeXT, Prompt=Please describe this image in detail2026.01 | 9.1 | — | 54.2 | 3.9 | — | — | — | |
| VCDBackbone=LLaVA-1.5 (7B), Prompt=Please describe this image in detail2026.01 | 9 | — | 42.9 | 4.6 | — | — | — | |
| M3IDBackbone=LLaVA-1.5 (7B), Prompt=Please describe this image in detail2026.01 | 9 | — | 40 | 3 | — | — | — | |
| Beam5Base Model=LLaVA-1.5, Decoding Strategy=Beam Search, Beam Size=52025.04 | 8.9 | 48.8 | 38.1 | 4.8 | — | — | — | |
| DoLa(high)Base Model=LLaVA-1.5, Decoding Strategy=DoLa, Premature Layer=high2025.04 | 8.8 | 52.2 | 40 | 4.2 | — | — | — | |
| OPERABackbone=LLaVA-1.5 (13B), Prompt=Please describe this image in detail2026.01 | 8.7 | — | 40 | 4.1 | — | — | — | |
| M3IDBackbone=LLaVA-NeXT, Prompt=Please describe this image in detail2026.01 | 8.7 | — | 51.9 | 3.1 | — | — | — | |
| VCDBase Model=LLaVA-1.5, Method Category=Contrastive Decoding2025.04 | 8.7 | 51.5 | 41.1 | 4.4 | — | — | — | |
| ICDBackbone=LLaVA-1.5 (13B), Prompt=Please describe this image in detail2026.01 | 8.5 | — | 39.5 | 4.1 | — | — | — | |
| VCDBackbone=LLaVA-1.5 (13B), Prompt=Please describe this image in detail2026.01 | 8.4 | — | 38.3 | 3.9 | — | — | — | |
| M3IDBackbone=LLaVA-1.5 (13B), Prompt=Please describe this image in detail2026.01 | 7.9 | — | 40 | 2.9 | — | — | — | |
| Beam3Base Model=LLaVA-1.5, Decoding Strategy=Beam Search, Beam Size=32025.04 | 7.9 | 49.7 | 37.5 | 4.6 | — | — | — | |
| LLaVA-v1.5-7BBase VLM Model=LLaVA-v1.5-7B2025.05 | 7.7 | 49.8 | 31.9 | 3.7 | — | — | — | |
| GreedyBase Model=LLaVA-1.5, Decoding Strategy=Greedy2025.04 | 7.6 | 49.5 | 32.1 | 3.8 | — | — | — | |
| DoLa(low)Base Model=LLaVA-1.5, Decoding Strategy=DoLa, Premature Layer=low2025.04 | 7.4 | 50.7 | 33.3 | 3.9 | — | — | — | |
| AGLABase Model=LLaVA-1.5, Method Category=Contrastive Decoding2025.04 | 7.4 | 51.1 | 34.6 | 3.9 | — | — | — | |
| CRoPSBackbone=LLaVA-NeXT, Prompt=Please describe this image in detail2026.01 | 7.2 | — | 44.6 | 2.6 | — | — | — | |
| SIDBackbone=LLaVA-1.5 (13B), Prompt=Please describe this image in detail2026.01 | 6.9 | — | 35 | 3.5 | — | — | — | |
| DPOBase VLM Model=LLaVA-v1.5-7B, Preference Learning Method=DPO2025.05 | 6.7 | 53.2 | 33.7 | 3.3 | — | — | — | |
| HA-DPOBase VLM Model=LLaVA-v1.5-7B, Preference Learning Method=HA-DPO2025.05 | 6.7 | 49.8 | 30.9 | 3.3 | — | — | — | |
| HALVABase VLM Model=LLaVA-v1.5-7B, Preference Learning Method=HALVA2025.05 | 6.6 | 53 | 32.2 | 3.4 | — | — | — | |
| V-DPOBase VLM Model=LLaVA-v1.5-7B, Preference Learning Method=V-DPO2025.05 | 6.6 | 49.1 | 30.8 | 3.1 | — | — | — | |
| SamplingBackbone=Qwen2-VL, Prompt=Please describe this image in detail2026.01 | 6.4 | — | 38.9 | 2.8 | — | — | — | |
| OPERABase Model=LLaVA-1.5, Decoding Strategy=Backtracking2025.04 | 6.4 | 49.7 | 29.1 | 2.9 | — | — | — | |
| LLaVA-v1.5-13BBase VLM Model=LLaVA-v1.5-13B2025.05 | 6.3 | 51 | 30.2 | 3 | — | — | — | |
| OPERABackbone=Qwen2-VL, Prompt=Please describe this image in detail2026.01 | 6.3 | — | 35 | 2.6 | — | — | — | |
| CRoPSBackbone=LLaVA-1.5 (7B), Prompt=Please describe this image in detail2026.01 | 6.3 | — | 29.3 | 2.8 | — | — | — | |
| DPOBase VLM Model=LLaVA-v1.5-13B, Preference Learning Method=DPO2025.05 | 6.2 | 54.3 | 31.8 | 2.6 | — | — | — | |
| ICDBackbone=Qwen2-VL, Prompt=Please describe this image in detail2026.01 | 6 | — | 34.1 | 2.5 | — | — | — | |
| ClearSightBackbone=Qwen2-VL, Prompt=Please describe this image in detail2026.01 | 5.9 | — | 34 | 2.2 | — | — | — | |
| VCDBackbone=Qwen2-VL, Prompt=Please describe this image in detail2026.01 | 5.7 | — | 33.9 | 2.4 | — | — | — | |
| CRoPSBackbone=LLaVA-1.5 (13B), Prompt=Please describe this image in detail2026.01 | 5.7 | — | 27.8 | 2.5 | — | — | — | |
| M3IDBackbone=Qwen2-VL, Prompt=Please describe this image in detail2026.01 | 5.5 | — | 27.9 | 1.5 | — | — | — | |
| SIDBackbone=Qwen2-VL, Prompt=Please describe this image in detail2026.01 | 5.4 | — | 30.6 | 1.8 | — | — | — | |
| CRoPSBackbone=Qwen2-VL, Prompt=Please describe this image in detail2026.01 | 5.1 | — | 24.2 | 1.1 | — | — | — | |
| mDPOBase VLM Model=LLaVA-v1.5-7B, Preference Learning Method=mDPO2025.05 | 5 | 52.5 | 27.5 | 2.4 | — | — | — | |
| TARACBase Model=LLaVA-1.5, Hyperparameters=alpha=0.3, beta=0.92025.04 | 5 | 48.3 | 27.1 | 2.5 | — | — | — | |
| mDPOBase VLM Model=LLaVA-v1.5-13B, Preference Learning Method=mDPO2025.05 | 4.6 | 52.6 | 25 | 2 | — | — | — | |
| LPOIBase VLM Model=LLaVA-v1.5-7B, Preference Learning Method=LPOI2025.05 | 4.3 | 51.9 | 26.4 | 2 | — | — | — | |
| LPOIBase VLM Model=LLaVA-v1.5-13B, Preference Learning Method=LPOI2025.05 | 3.9 | 52.9 | 22.3 | 1.8 | — | — | — | |
| DPOBase VLM Model=Idefics2-8B, Preference Learning Method=DPO2025.05 | 3.5 | 37.4 | 8.1 | 0.2 | — | — | — | |
| Idefics2-8BBase VLM Model=Idefics2-8B2025.05 | 3.4 | 36.5 | 7.6 | 0.4 | — | — | — | |
| mDPOBase VLM Model=Idefics2-8B, Preference Learning Method=mDPO2025.05 | 2.7 | 37.7 | 6.2 | 0.2 | — | — | — | |
| LPOIBase VLM Model=Idefics2-8B, Preference Learning Method=LPOI2025.05 | 2.6 | 36.4 | 5.7 | 0.2 | — | — | — | |
| GIFTModel=LLaVA-1.5 7B2025.10 | — | — | — | — | 87.2 | 80.9 | 6.6 | |
| GIFTModel=LLaVA-1.5 13B2025.10 | — | — | — | — | 85.7 | 76.8 | 5.5 | |
| GIFTModel=Qwen2-VL 7B2025.10 | — | — | — | — | 85.4 | 76.2 | 5.5 | |
| GIFTModel=Qwen3-VL 8B2025.10 | — | — | — | — | 83.6 | 74 | 6.8 | |
| GreedyModel=LLaVA-1.5 7B2025.10 | — | — | — | — | 83.9 | 74.9 | 7.2 | |
| GreedyModel=LLaVA-1.5 13B2025.10 | — | — | — | — | 83.4 | 73.5 | 6.7 | |
| GreedyModel=Qwen2-VL 7B2025.10 | — | — | — | — | 84.1 | 74.4 | 6.2 | |
| GreedyModel=Qwen3-VL 8B2025.10 | — | — | — | — | 80.2 | 68.3 | 7.9 |