Object Hallucination Evaluation on CHAIR MSCOCO
62CS ScoreGreedy
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| GreedyBackbone=Shikra-7B, Max new tokens=512, Category=Decoding Strategy2026.05 | 62 | 17.5 | 70.9 | |
| BeamBackbone=Shikra-7B, Max new tokens=512, Category=Decoding Strategy2026.05 | 59.2 | 16.2 | 74.2 | |
| VCDBackbone=LLaVA-1.5-13B, Max new tokens=512, Category=Contrastive Decoding, Decoding strategy=Greedy2026.05 | 57.2 | 16.1 | 71.3 | |
| VCDBackbone=Shikra-7B, Max new tokens=512, Category=Contrastive Decoding, Decoding strategy=Greedy2026.05 | 56.4 | 15.5 | 75.2 | |
| BeamBackbone=LLaVA-1.5-7B, Max new tokens=512, Category=Decoding Strategy2026.05 | 56.2 | 15.1 | 74 | |
| GreedyBackbone=LLaVA-1.5-7B, Max new tokens=512, Category=Decoding Strategy2026.05 | 55.4 | 14.4 | 74 | |
| VCDBackbone=LLaVA-1.5-7B, Max new tokens=512, Category=Contrastive Decoding, Decoding strategy=Greedy2026.05 | 54.6 | 17.3 | 70.1 | |
| MemVRBackbone=LLaVA-1.5-13B, Max new tokens=512, Category=Hidden states-intervention2026.05 | 53.2 | 13.5 | 75.6 | |
| VCDBackbone=Average, Max new tokens=512, Category=Contrastive Decoding, Decoding strategy=Greedy2026.05 | 51 | 15.9 | — | |
| MemVRBackbone=LLaVA-1.5-7B, Max new tokens=512, Category=Hidden states-intervention2026.05 | 50.4 | 14.3 | 74.2 | |
| BeamBackbone=LLaVA-1.5-13B, Max new tokens=512, Category=Decoding Strategy2026.05 | 50.4 | 15.2 | 76.3 | |
| GreedyBackbone=LLaVA-1.5-13B, Max new tokens=512, Category=Decoding Strategy2026.05 | 49.7 | 14.7 | 71 | |
| MemVRBackbone=Shikra-7B, Max new tokens=512, Category=Hidden states-intervention2026.05 | 48.8 | 16.5 | 70.9 | |
| OPERABackbone=LLaVA-1.5-7B, Max new tokens=512, Category=Decoding Strategy2026.05 | 48.2 | 13.1 | 76.9 | |
| BeamBackbone=Average, Max new tokens=512, Category=Decoding Strategy2026.05 | 47 | 13.9 | — | |
| GreedyBackbone=Average, Max new tokens=512, Category=Decoding Strategy2026.05 | 46.9 | 13.3 | — | |
| VCDBackbone=Qwen-VL, Max new tokens=512, Category=Contrastive Decoding, Decoding strategy=Greedy2026.05 | 45.4 | 18.1 | 70 | |
| MemVRBackbone=Average, Max new tokens=512, Category=Hidden states-intervention2026.05 | 44.8 | 13.5 | — | |
| OPERABackbone=Shikra-7B, Max new tokens=512, Category=Decoding Strategy2026.05 | 41.9 | 13.8 | 72.1 | |
| VCDBackbone=MiniGPT-4-7B, Max new tokens=512, Category=Contrastive Decoding, Decoding strategy=Greedy2026.05 | 41.4 | 12.6 | 68.2 | |
| OPERABackbone=LLaVA-1.5-13B, Max new tokens=512, Category=Decoding Strategy2026.05 | 41.3 | 14.1 | 77.2 | |
| GreedyBackbone=MiniGPT-4-7B, Max new tokens=512, Category=Decoding Strategy2026.05 | 39.4 | 11 | 80.5 | |
| BeamBackbone=MiniGPT-4-7B, Max new tokens=512, Category=Decoding Strategy2026.05 | 39.2 | 12.2 | 81.7 | |
| OPERABackbone=Average, Max new tokens=512, Category=Decoding Strategy2026.05 | 38.4 | 12.3 | — | |
| MemVRBackbone=Qwen-VL, Max new tokens=512, Category=Hidden states-intervention2026.05 | 38.2 | 13.6 | 74.9 | |
| MemVRBackbone=MiniGPT-4-7B, Max new tokens=512, Category=Hidden states-intervention2026.05 | 33.6 | 9.7 | 83.5 | |
| HGAIBackbone=Shikra-7B, Max new tokens=512, Category=Attention-intervention, Decoding strategy=Greedy2026.05 | 32.8 | 10.4 | 81 | |
| OPERABackbone=MiniGPT-4-7B, Max new tokens=512, Category=Decoding Strategy2026.05 | 30.9 | 11.2 | 77.5 | |
| BeamBackbone=Qwen-VL, Max new tokens=512, Category=Decoding Strategy2026.05 | 30 | 10.7 | 83.8 | |
| OPERABackbone=Qwen-VL, Max new tokens=512, Category=Decoding Strategy2026.05 | 29.6 | 9.5 | 79.4 | |
| GreedyBackbone=Qwen-VL, Max new tokens=512, Category=Decoding Strategy2026.05 | 28.2 | 8.9 | 84.9 | |
| HGAIBackbone=MiniGPT-4-7B, Max new tokens=512, Category=Attention-intervention, Decoding strategy=Greedy2026.05 | 28 | 7.5 | 85.1 | |
| HGAIBackbone=LLaVA-1.5-13B, Max new tokens=512, Category=Attention-intervention, Decoding strategy=Greedy2026.05 | 27.1 | 7.9 | 85.1 | |
| HGAIBackbone=Average, Max new tokens=512, Category=Attention-intervention, Decoding strategy=Greedy2026.05 | 26.7 | 7.5 | — | |
| HGAIBackbone=LLaVA-1.5-7B, Max new tokens=512, Category=Attention-intervention, Decoding strategy=Greedy2026.05 | 24.6 | 6.4 | 86.8 | |
| HGAIBackbone=Qwen-VL, Max new tokens=512, Category=Attention-intervention, Decoding strategy=Greedy2026.05 | 20.9 | 5.2 | 89.1 | |
| AFIPBackbone=LLaVA-1.5-13B, Max new tokens=512, Category=Attention-intervention, Decoding strategy=Greedy2026.05 | 18.4 | 6.4 | 89.2 | |
| AFIPBackbone=Shikra-7B, Max new tokens=512, Category=Attention-intervention, Decoding strategy=Greedy2026.05 | 18 | 6.6 | 89.8 | |
| AFIPBackbone=MiniGPT-4-7B, Max new tokens=512, Category=Attention-intervention, Decoding strategy=Greedy2026.05 | 17.8 | 7.3 | 89.5 | |
| AFIPBackbone=Average, Max new tokens=512, Category=Attention-intervention, Decoding strategy=Greedy2026.05 | 17.5 | 5.9 | — | |
| AFIPBackbone=LLaVA-1.5-7B, Max new tokens=512, Category=Attention-intervention, Decoding strategy=Greedy2026.05 | 16.8 | 4.4 | 91.3 | |
| AFIPBackbone=Qwen-VL, Max new tokens=512, Category=Attention-intervention, Decoding strategy=Greedy2026.05 | 16.6 | 5 | 91.1 |