Hallucination Detection on TyDiQA (test)
88.4AUROCHARP
Evaluation Results
| Method | Links | |
|---|---|---|
| HARPModels=Qwen-2.5-7B-Instruct, Single=true2025.09 | 88.4 | |
| HARPModels=LLaMA-3.1-8B, Single=true2025.09 | 86.6 | |
| EigenScoreModels=LLaMA-3.1-8B, Single=false2025.09 | 82.4 | |
| EigenScoreModels=Qwen-2.5-7B-Instruct, Single=false2025.09 | 74.8 | |
| Lexical SimilarityModels=LLaMA-3.1-8B, Single=false2025.09 | 69.5 | |
| HaloScopeModels=Qwen-2.5-7B-Instruct, Single=true2025.09 | 69 | |
| Semantic EntropyModels=Qwen-2.5-7B-Instruct, Single=false2025.09 | 68.6 | |
| Semantic EntropyModels=LLaMA-3.1-8B, Single=false2025.09 | 62.2 | |
| Lexical SimilarityModels=Qwen-2.5-7B-Instruct, Single=false2025.09 | 60.3 | |
| PerplexityModels=LLaMA-3.1-8B, Single=true2025.09 | 53.4 | |
| HaloScopeModels=LLaMA-3.1-8B, Single=true2025.09 | 53.3 | |
| LN-EntropyModels=LLaMA-3.1-8B, Single=false2025.09 | 48.8 | |
| LN-EntropyModels=Qwen-2.5-7B-Instruct, Single=false2025.09 | 47.1 | |
| PerplexityModels=Qwen-2.5-7B-Instruct, Single=true2025.09 | 30.5 |