Hallucination Detection on CSQA
85.1AUROCOSCAR
Evaluation Results
| Method | Links | |
|---|---|---|
| OSCARBackbone=Dream-7B-Instruct, Method Category=Traj., Sample Count=64, Evaluation Protocol=LLM-as-Judge2026.04 | 85.1 | |
| TraceDetBackbone=Dream-7B-Instruct, Method Category=Traj., Sample Count=128, Evaluation Protocol=LLM-as-Judge, Requires trained classifier=true2026.04 | 84.7 | |
| DynHDBackbone=Dream-7B-Instruct, Method Category=Traj., Sample Count=64, Evaluation Protocol=LLM-as-Judge, Requires trained classifier=true2026.04 | 84.6 | |
| OSCARBackbone=Dream-7B-Instruct, Method Category=Traj., Sample Count=128, Evaluation Protocol=LLM-as-Judge2026.04 | 84.3 | |
| OSCARBackbone=Dream-7B-Instruct, Method Category=Traj., Sample Count=128, Evaluation Protocol=EM2026.04 | 84.2 | |
| TraceDetBackbone=Dream-7B-Instruct, Method Category=Traj., Sample Count=64, Evaluation Protocol=LLM-as-Judge, Requires trained classifier=true2026.04 | 84.1 | |
| OSCARBackbone=Dream-7B-Instruct, Method Category=Traj., Sample Count=64, Evaluation Protocol=EM2026.04 | 83.7 | |
| DynHDBackbone=Dream-7B-Instruct, Method Category=Traj., Sample Count=128, Evaluation Protocol=LLM-as-Judge, Requires trained classifier=true2026.04 | 83.5 | |
| OSCARBackbone=LLaDA-8B-Instruct, Method Category=Traj., Sample Count=128, Evaluation Protocol=LLM-as-Judge2026.04 | 83.4 | |
| OSCARBackbone=LLaDA-8B-Instruct, Method Category=Traj., Sample Count=64, Evaluation Protocol=LLM-as-Judge2026.04 | 82.8 | |
| DynHDBackbone=LLaDA-8B-Instruct, Method Category=Traj., Sample Count=128, Evaluation Protocol=LLM-as-Judge, Requires trained classifier=true2026.04 | 81.6 | |
| DynHDBackbone=LLaDA-8B-Instruct, Method Category=Traj., Sample Count=64, Evaluation Protocol=LLM-as-Judge, Requires trained classifier=true2026.04 | 81.3 | |
| Mistral-7B-Instruct-v0.3Noise-Enhanced Sampling=true, K (Number of samples)=102025.02 | 79.55 | |
| OSCARBackbone=LLaDA-8B-Instruct, Method Category=Traj., Sample Count=128, Evaluation Protocol=EM2026.04 | 79.4 | |
| OSCARBackbone=LLaDA-8B-Instruct, Method Category=Traj., Sample Count=64, Evaluation Protocol=EM2026.04 | 78.9 | |
| EigenScoreBackbone=Dream-7B-Instruct, Method Category=Lat., Sample Count=642026.04 | 77.5 | |
| Lexical SimilarityBackbone=Dream-7B-Instruct, Method Category=Out., Sample Count=1282026.04 | 77.3 | |
| TraceDetBackbone=LLaDA-8B-Instruct, Method Category=Traj., Sample Count=128, Evaluation Protocol=LLM-as-Judge, Requires trained classifier=true2026.04 | 77.2 | |
| TraceDetBackbone=LLaDA-8B-Instruct, Method Category=Traj., Sample Count=64, Evaluation Protocol=LLM-as-Judge, Requires trained classifier=true2026.04 | 77.1 | |
| Lexical SimilarityBackbone=Dream-7B-Instruct, Method Category=Out., Sample Count=642026.04 | 76.9 | |
| EigenScoreBackbone=Dream-7B-Instruct, Method Category=Lat., Sample Count=1282026.04 | 76.9 | |
| Phi-3-mini-4k-instruct (3.8B)Noise-Enhanced Sampling=true, K (Number of samples)=102025.02 | 76.6 | |
| Mistral-7B-Instruct-v0.3Noise-Enhanced Sampling=false, K (Number of samples)=102025.02 | 76.52 | |
| Phi-3-mini-4k-instruct (3.8B)Noise-Enhanced Sampling=false, K (Number of samples)=102025.02 | 75.05 | |
| Llama-3.2-3B-InstructNoise-Enhanced Sampling=true, K (Number of samples)=102025.02 | 72.83 | |
| fDBDModel=Qwen-2.5-7B-Instruct, Single Sample=true, k=selected (100)2026.02 | 72.47 | |
| TDGNetModel=Dream-7B-Instruct, Approach Category=Latent-Based, SS (Single Sampling)=true2026.02 | 72 | |
| NCIModel=Qwen-2.5-7B-Instruct, Single Sample=true2026.02 | 71.6 | |
| Llama-2-7B-chatNoise-Enhanced Sampling=true, K (Number of samples)=102025.02 | 71.56 | |
| fDBDModel=Qwen-2.5-7B-Instruct, Single Sample=true, k=ALL2026.02 | 71.5 | |
| TSVModel=Dream-7B-Instruct, Approach Category=Latent-Based, SS (Single Sampling)=true2026.02 | 71 | |
| Llama-3.2-3B-InstructNoise-Enhanced Sampling=false, K (Number of samples)=102025.02 | 70.72 | |
| Llama-2-7B-chatNoise-Enhanced Sampling=false, K (Number of samples)=102025.02 | 70.59 | |
| fDBDModel=Llama-3.2-3B-Instruct, Single Sample=true, k=selected (1000)2026.02 | 69.24 | |
| Llama-2-13B-chatNoise-Enhanced Sampling=true, K (Number of samples)=102025.02 | 69.1 | |
| fDBDModel=Llama-3.2-3B-Instruct, Single Sample=true, k=ALL2026.02 | 68.15 | |
| P(True)Model=Qwen-2.5-7B-Instruct, Single Sample=true2026.02 | 68.01 | |
| Lexical SimilarityModel=Dream-7B-Instruct, Approach Category=Output-Based, SS (Single Sampling)=false2026.02 | 68 | |
| Llama-2-13B-chatNoise-Enhanced Sampling=false, K (Number of samples)=102025.02 | 67.55 | |
| CoE-CModel=Qwen-2.5-7B-Instruct, Single Sample=true2026.02 | 66.89 | |
| NCIModel=Llama-3.2-3B-Instruct, Single Sample=true2026.02 | 66.07 | |
| Max PModel=Llama-3.2-3B-Instruct, Single Sample=true2026.02 | 66.01 | |
| fDBDBackbone=Qwen-3-32B, hyperparameter k=all k2026.02 | 65.68 | |
| PerplexityBackbone=LLaDA-8B-Instruct, Method Category=Output, Sample Count=1282026.04 | 65.6 | |
| LN Predictive ProbabilityModel=Qwen-2.5-7B-Instruct, Single Sample=true2026.02 | 65.19 | |
| TDGNetSource=CSQA, Zero-shot=true, Base Model=LLaDA2026.02 | 65 | |
| TDGNetModel=LLaDA-8B-Instruct, Approach Category=Latent-Based, SS (Single Sampling)=true2026.02 | 65 | |
| PerplexityBackbone=LLaDA-8B-Instruct, Method Category=Output, Sample Count=642026.04 | 65 | |
| Predictive ProbabilityModel=Qwen-2.5-7B-Instruct, Single Sample=true2026.02 | 64.91 | |
| LN-EntropyBackbone=LLaDA-8B-Instruct, Method Category=Output, Sample Count=1282026.04 | 64.6 | |
| NCIBackbone=Qwen-3-32B2026.02 | 64.54 | |
| LN-EntropyBackbone=LLaDA-8B-Instruct, Method Category=Output, Sample Count=642026.04 | 64.4 | |
| SelfCheckGPT NLIModel=Llama-3.2-3B-Instruct, Single Sample=false2026.02 | 64.18 | |
| Semantic EntropyModel=LLaDA-8B-Instruct, Approach Category=Output-Based, SS (Single Sampling)=false2026.02 | 64 | |
| PerplexityModel=Llama-3.2-3B-Instruct, Single Sample=true2026.02 | 63.23 | |
| Lexical SimilarityModel=Llama-3.2-3B-Instruct, Single Sample=false2026.02 | 62.94 | |
| CoE-RModel=Qwen-2.5-7B-Instruct, Single Sample=true2026.02 | 62.75 | |
| TSVBackbone=Dream-7B-Instruct, Method Category=Lat., Sample Count=1282026.04 | 62.3 | |
| PerplexityModel=Qwen-2.5-7B-Instruct, Single Sample=true2026.02 | 61.94 | |
| Gemma-2B-itNoise-Enhanced Sampling=true, K (Number of samples)=102025.02 | 61.71 | |
| Predictive ProbabilityModel=Llama-3.2-3B-Instruct, Single Sample=true2026.02 | 61.63 | |
| LN Predictive ProbabilityModel=Llama-3.2-3B-Instruct, Single Sample=true2026.02 | 61.51 | |
| TSVModel=LLaDA-8B-Instruct, Approach Category=Latent-Based, SS (Single Sampling)=true2026.02 | 61 | |
| Lexical SimilarityBackbone=LLaDA-8B-Instruct, Method Category=Output, Sample Count=642026.04 | 60.7 | |
| Semantic EntropyModel=Llama-3.2-3B-Instruct, Single Sample=false2026.02 | 60.61 | |
| EigenScoreBackbone=LLaDA-8B-Instruct, Method Category=Latent, Sample Count=642026.04 | 60.6 | |
| Lexical SimilarityModel=Qwen-2.5-7B-Instruct, Single Sample=false2026.02 | 60.57 | |
| SelfCheckGPT NLIModel=Qwen-2.5-7B-Instruct, Single Sample=false2026.02 | 60.18 | |
| PerplexityModel=LLaDA-8B-Instruct, Approach Category=Output-Based, SS (Single Sampling)=false2026.02 | 60 | |
| LN Predictive ProbabilityBackbone=Qwen-3-32B2026.02 | 59.35 | |
| Predictive ProbabilityBackbone=Qwen-3-32B2026.02 | 59.3 | |
| Semantic EntropyModel=Qwen-2.5-7B-Instruct, Single Sample=false2026.02 | 59.1 | |
| TDGNetSource=Hotpot, Zero-shot=true, Base Model=LLaDA2026.02 | 59 | |
| LN-EntropyModel=LLaDA-8B-Instruct, Approach Category=Output-Based, SS (Single Sampling)=false2026.02 | 59 | |
| LN-EntropyModel=Dream-7B-Instruct, Approach Category=Output-Based, SS (Single Sampling)=false2026.02 | 59 | |
| Gemma-2B-itNoise-Enhanced Sampling=false, K (Number of samples)=102025.02 | 58.97 | |
| CoE-CModel=Llama-3.2-3B-Instruct, Single Sample=true2026.02 | 58.82 | |
| EigenScoreBackbone=LLaDA-8B-Instruct, Method Category=Latent, Sample Count=1282026.04 | 58.5 | |
| CCSBackbone=LLaDA-8B-Instruct, Method Category=Latent, Sample Count=642026.04 | 58.5 | |
| PerplexityBackbone=Qwen-3-32B2026.02 | 58.44 | |
| Lexical SimilarityBackbone=LLaDA-8B-Instruct, Method Category=Output, Sample Count=1282026.04 | 57.3 | |
| PerplexityModel=Dream-7B-Instruct, Approach Category=Output-Based, SS (Single Sampling)=false2026.02 | 57 | |
| TSVBackbone=Dream-7B-Instruct, Method Category=Lat., Sample Count=642026.04 | 56.8 | |
| CCSSource=CSQA, Zero-shot=true, Base Model=LLaDA2026.02 | 56 | |
| CCSModel=LLaDA-8B-Instruct, Approach Category=Latent-Based, SS (Single Sampling)=true2026.02 | 56 | |
| CCSModel=Dream-7B-Instruct, Approach Category=Latent-Based, SS (Single Sampling)=true2026.02 | 56 | |
| TSVBackbone=LLaDA-8B-Instruct, Method Category=Latent, Sample Count=642026.04 | 55.2 | |
| CCSSource=Hotpot, Zero-shot=true, Base Model=LLaDA2026.02 | 55 | |
| TDGNetSource=Trivia, Zero-shot=true, Base Model=LLaDA2026.02 | 55 | |
| EigenScoreModel=Dream-7B-Instruct, Approach Category=Latent-Based, SS (Single Sampling)=false2026.02 | 55 | |
| CCSBackbone=Dream-7B-Instruct, Method Category=Lat., Sample Count=1282026.04 | 54.2 | |
| EigenScoreModel=LLaDA-8B-Instruct, Approach Category=Latent-Based, SS (Single Sampling)=false2026.02 | 54 | |
| CCSBackbone=Dream-7B-Instruct, Method Category=Lat., Sample Count=642026.04 | 53.2 | |
| Lexical SimilarityModel=LLaDA-8B-Instruct, Approach Category=Output-Based, SS (Single Sampling)=false2026.02 | 53 | |
| TSVBackbone=LLaDA-8B-Instruct, Method Category=Latent, Sample Count=1282026.04 | 52.9 | |
| CCSSource=Math, Zero-shot=true, Base Model=LLaDA2026.02 | 52 | |
| Semantic EntropyBackbone=Dream-7B-Instruct, Method Category=Out., Sample Count=1282026.04 | 51.4 | |
| CCSSource=Trivia, Zero-shot=true, Base Model=LLaDA2026.02 | 51 | |
| CCSBackbone=LLaDA-8B-Instruct, Method Category=Latent, Sample Count=1282026.04 | 50.5 | |
| Semantic EntropyModel=Dream-7B-Instruct, Approach Category=Output-Based, SS (Single Sampling)=false2026.02 | 50 |