Object Hallucination Evaluation on POPE (average across random and popular)
91.56Accuracy (POPE)R-CoV
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| R-CoVBase Model=LLaVA-1.52026.04 | 91.56 | — | 91.25 | |
| LogicCheckGPTBase Model=LLaVA-1.52026.04 | 91 | — | 90.81 | |
| SelfCheckBase Model=LLaVA-1.52026.04 | 89.22 | — | 88.8 | |
| R-CoVBase Model=Qwen2.5-VL2026.04 | 88 | — | 86.83 | |
| LUREBase Model=LLaVA-1.52026.04 | 87.33 | — | 87.67 | |
| R-CoVBase Model=mPLUG-Owl2026.04 | 87 | — | 87 | |
| VanillaBase Model=LLaVA-1.52026.04 | 87 | — | 87.92 | |
| LogicCheckGPTBase Model=Qwen2.5-VL2026.04 | 87 | — | 85.39 | |
| VanillaBase Model=Qwen2.5-VL2026.04 | 86.22 | — | 84.19 | |
| R-CoVBase Model=MiniGPT-42026.04 | 85.89 | — | 84.79 | |
| VISTABase Model=DeepSeek-VL-Chat, Decoding Strategy=nucleus sampling, Max New Tokens=322026.04 | 85.88 | — | 84.51 | |
| PTIBase Model=Qwen-VL-Chat, Decoding Strategy=nucleus sampling, Max New Tokens=322026.04 | 85.69 | — | 84.62 | |
| PAIBase Model=DeepSeek-VL-Chat, Decoding Strategy=nucleus sampling, Max New Tokens=322026.04 | 85.57 | — | 84.82 | |
| VCDBase Model=DeepSeek-VL-Chat, Decoding Strategy=nucleus sampling, Max New Tokens=322026.04 | 85.45 | — | 84.47 | |
| PTIBase Model=DeepSeek-VL-Chat, Decoding Strategy=nucleus sampling, Max New Tokens=322026.04 | 85.14 | — | 85.01 | |
| VCDBase Model=Qwen-VL-Chat, Decoding Strategy=nucleus sampling, Max New Tokens=322026.04 | 85.06 | — | 83.72 | |
| LogicCheckGPTBase Model=mPLUG-Owl2026.04 | 85 | — | 84.84 | |
| PAIBase Model=Qwen-VL-Chat, Decoding Strategy=nucleus sampling, Max New Tokens=322026.04 | 84.91 | — | 84.15 | |
| VTIBase Model=DeepSeek-VL-Chat, Decoding Strategy=nucleus sampling, Max New Tokens=322026.04 | 84.81 | — | 84.65 | |
| VanillaBase Model=DeepSeek-VL-Chat, Decoding Strategy=nucleus sampling, Max New Tokens=322026.04 | 83.89 | — | 83.06 | |
| VanillaBase Model=Qwen-VL-Chat, Decoding Strategy=nucleus sampling, Max New Tokens=322026.04 | 83.69 | — | 82.92 | |
| VISTABase Model=Qwen-VL-Chat, Decoding Strategy=nucleus sampling, Max New Tokens=322026.04 | 83.54 | — | 82.12 | |
| VTIBase Model=Qwen-VL-Chat, Decoding Strategy=nucleus sampling, Max New Tokens=322026.04 | 83.27 | — | 82.19 | |
| LogicCheckGPTBase Model=MiniGPT-42026.04 | 82.67 | — | 80.67 | |
| PTIBase Model=LLAVA-1.5, Decoding Strategy=nucleus sampling, Max New Tokens=322026.04 | 82.21 | — | 82.85 | |
| PAIBase Model=LLAVA-1.5, Decoding Strategy=nucleus sampling, Max New Tokens=322026.04 | 81.73 | — | 82.95 | |
| VCDBase Model=LLAVA-1.5, Decoding Strategy=nucleus sampling, Max New Tokens=322026.04 | 81.53 | — | 82.8 | |
| LUREBase Model=MiniGPT-42026.04 | 80.67 | — | 81.29 | |
| VTIBase Model=LLAVA-1.5, Decoding Strategy=nucleus sampling, Max New Tokens=322026.04 | 80.41 | — | 81.64 | |
| VISTABase Model=LLAVA-1.5, Decoding Strategy=nucleus sampling, Max New Tokens=322026.04 | 80.08 | — | 81.98 | |
| VanillaBase Model=LLAVA-1.5, Decoding Strategy=nucleus sampling, Max New Tokens=322026.04 | 79.88 | — | 81.23 | |
| LUREBase Model=mPLUG-Owl2026.04 | 79.11 | — | 78.81 | |
| LRV-InstructionBase Model=MiniGPT-42026.04 | 78.67 | — | 76.93 | |
| VanillaBase Model=MiniGPT-42026.04 | 78.44 | — | 80.11 | |
| SelfCheckBase Model=MiniGPT-42026.04 | 75.22 | — | 74.04 | |
| SelfCheckBase Model=mPLUG-Owl2026.04 | 69.78 | — | 75.73 | |
| LRV-InstructionBase Model=mPLUG-Owl2026.04 | 67.44 | — | 73.75 | |
| VanillaBase Model=mPLUG-Owl2026.04 | 52.56 | — | 67.68 | |
| BaselineModel=LLaVA-1.5-Qwen2.5-7B, Training Protocol=LoRA2026.04 | — | 87.7 | — | |
| V-GIFTModel=LLaVA-1.5-Qwen2.5-7B, Training Protocol=LoRA2026.04 | — | 88.5 | — | |
| VIRALModel=LLaVA-1.5-Qwen2.5-7B, Training Protocol=LoRA2026.04 | — | 88.3 | — |