Hallucination Detection on AQuA
0.7822AUROCfDBD
Evaluation Results
| Method | Links | |
|---|---|---|
| fDBDModel=Qwen-2.5-7B-Instruct, Single Sample=true, k=selected (100)2026.02 | 0.7822 | |
| NCIModel=Qwen-2.5-7B-Instruct, Single Sample=true2026.02 | 0.7819 | |
| fDBDModel=Qwen-2.5-7B-Instruct, Single Sample=true, k=ALL2026.02 | 0.7708 | |
| fDBDModel=Llama-3.2-3B-Instruct, Single Sample=true, k=selected (100)2026.02 | 0.762 | |
| fDBDModel=Llama-3.2-3B-Instruct, Single Sample=true, k=ALL2026.02 | 0.758 | |
| NCIModel=Llama-3.2-3B-Instruct, Single Sample=true2026.02 | 0.7441 | |
| LN Predictive ProbabilityModel=Qwen-2.5-7B-Instruct, Single Sample=true2026.02 | 0.7417 | |
| Predictive ProbabilityModel=Qwen-2.5-7B-Instruct, Single Sample=true2026.02 | 0.7337 | |
| P(True)Model=Qwen-2.5-7B-Instruct, Single Sample=true2026.02 | 0.7286 | |
| PerplexityModel=Llama-3.2-3B-Instruct, Single Sample=true2026.02 | 0.7285 | |
| Lexical SimilarityModel=Qwen-2.5-7B-Instruct, Single Sample=false2026.02 | 0.7262 | |
| CoE-RModel=Qwen-2.5-7B-Instruct, Single Sample=true2026.02 | 0.7213 | |
| CoE-CModel=Qwen-2.5-7B-Instruct, Single Sample=true2026.02 | 0.7204 | |
| PerplexityModel=Qwen-2.5-7B-Instruct, Single Sample=true2026.02 | 0.7166 | |
| Lexical SimilarityModel=Llama-3.2-3B-Instruct, Single Sample=false2026.02 | 0.7148 | |
| SelfCheckGPT NLIModel=Qwen-2.5-7B-Instruct, Single Sample=false2026.02 | 0.709 | |
| Semantic EntropyModel=Qwen-2.5-7B-Instruct, Single Sample=false2026.02 | 0.6962 | |
| Predictive ProbabilityModel=Llama-3.2-3B-Instruct, Single Sample=true2026.02 | 0.6907 | |
| LN Predictive ProbabilityModel=Llama-3.2-3B-Instruct, Single Sample=true2026.02 | 0.6898 | |
| Max PModel=Llama-3.2-3B-Instruct, Single Sample=true2026.02 | 0.6602 | |
| SelfCheckGPT NLIModel=Llama-3.2-3B-Instruct, Single Sample=false2026.02 | 0.6601 | |
| fDBDBackbone=Qwen-3-32B, hyperparameter k=all k2026.02 | 0.6537 | |
| Semantic EntropyModel=Llama-3.2-3B-Instruct, Single Sample=false2026.02 | 0.6471 | |
| CoE-CModel=Llama-3.2-3B-Instruct, Single Sample=true2026.02 | 0.6256 | |
| NCIBackbone=Qwen-3-32B2026.02 | 0.6028 | |
| LN Predictive ProbabilityBackbone=Qwen-3-32B2026.02 | 0.5821 | |
| Predictive ProbabilityBackbone=Qwen-3-32B2026.02 | 0.5797 | |
| PerplexityBackbone=Qwen-3-32B2026.02 | 0.5389 | |
| Max PModel=Qwen-2.5-7B-Instruct, Single Sample=true2026.02 | 0.5083 | |
| CoE-RModel=Llama-3.2-3B-Instruct, Single Sample=true2026.02 | 0.4555 | |
| P(True)Model=Llama-3.2-3B-Instruct, Single Sample=true2026.02 | 0.3938 |