Hallucination Detection on GSM8K
90.37AUROCARS (CCS)
Evaluation Results
| Method | Links | |
|---|---|---|
| ARS (CCS)Model=Qwen3-8B, Single Sampling=true, Supervision=false2026.01 | 90.37 | |
| ARS (Probing)Model=Qwen3-8B, Single Sampling=true, Supervision=true2026.01 | 89.88 | |
| SinkProbeLLM=Phi3.52026.04 | 85.4 | |
| SinkProbeLLM=Llama3.2-3B2026.04 | 84.5 | |
| MTopDivLLM=Phi3.52026.04 | 84.5 | |
| LookbackLensLLM=Phi3.52026.04 | 84.3 | |
| LookbackLensLLM=Mistral-Nemo2026.04 | 84.1 | |
| LSModel=Llama-3.2-3B-Instruct2025.08 | 83.79 | |
| G-DetectorModel=Qwen3-8B, Single Sampling=true, Supervision=true2026.01 | 83.78 | |
| LookbackLensLLM=Llama3.2-3B2026.04 | 83.5 | |
| TSVModel=Qwen3-8B, Single Sampling=true, Supervision=true2026.01 | 83.15 | |
| Llama-3.2-3B-InstructNoise-Enhanced Sampling=true, K (Number of samples)=102025.02 | 82.7 | |
| LapEigvalLLM=Llama3.1-8B2026.04 | 82.6 | |
| LapEigvalLLM=Llama3.2-3B2026.04 | 82.5 | |
| SinkProbeLLM=Llama3.1-8B2026.04 | 82.4 | |
| AttnLogDetLLM=Llama3.2-3B2026.04 | 81.9 | |
| MTopDivLLM=Llama3.2-3B2026.04 | 81.6 | |
| LookbackLensLLM=Llama3.1-8B2026.04 | 81.6 | |
| MTopDivLLM=Mistral-Nemo2026.04 | 81.3 | |
| AttnLogDetLLM=Llama3.1-8B2026.04 | 81.2 | |
| SinkProbeLLM=Mistral-Nemo2026.04 | 81.1 | |
| fDBDBackbone=Qwen-3-32B, hyperparameter k=all k2026.02 | 80.6 | |
| LapEigvalLLM=Mistral-Nemo2026.04 | 80.4 | |
| MTopDivLLM=Llama3.1-8B2026.04 | 80.3 | |
| AttnEigvalLLM=Llama3.2-3B2026.04 | 80.2 | |
| AttnEigvalLLM=Llama3.1-8B2026.04 | 79.7 | |
| AttnLogDetLLM=Mistral-Nemo2026.04 | 79.7 | |
| AttnEigvalLLM=Phi3.52026.04 | 79.4 | |
| Llama-2-13B-chatNoise-Enhanced Sampling=true, K (Number of samples)=102025.02 | 79.25 | |
| Mistral-7B-Instruct-v0.3Noise-Enhanced Sampling=true, K (Number of samples)=102025.02 | 78.5 | |
| TSVModel=DeepSeek-R1-Distill-Llama-8B, Single Sampling=true, Supervision=true2026.01 | 78.29 | |
| AttnLogDetLLM=Phi3.52026.04 | 78.2 | |
| AttnEigvalLLM=Mistral-Nemo2026.04 | 78 | |
| ARS (Probing)Model=DeepSeek-R1-Distill-Llama-8B, Single Sampling=true, Supervision=true2026.01 | 77.62 | |
| LapEigvalLLM=Phi3.52026.04 | 77.5 | |
| Llama-2-13B-chatNoise-Enhanced Sampling=false, K (Number of samples)=102025.02 | 77.2 | |
| fDBDModel=Qwen-2.5-7B-Instruct, Single Sample=true, k=ALL2026.02 | 77.19 | |
| fDBDModel=Qwen-2.5-7B-Instruct, Single Sample=true, k=selected (All)2026.02 | 77.19 | |
| PerplexityBackbone=Qwen-3-32B2026.02 | 76.86 | |
| FAVAModel=gpt-4o-mini-2024-07-18, Category=Multi-trajectory / LLM-based2026.05 | 76.6 | |
| Llama-3.2-3B-InstructNoise-Enhanced Sampling=false, K (Number of samples)=102025.02 | 76.53 | |
| fDBDModel=Llama-3.2-3B-Instruct, Single Sample=true, k=selected (100)2026.02 | 76.36 | |
| NCIModel=Llama-3.2-3B-Instruct, Single Sample=true2026.02 | 76.32 | |
| SelfCheckGPT NLIModel=Qwen-2.5-7B-Instruct, Single Sample=false2026.02 | 76.22 | |
| Llama-2-7B-chatNoise-Enhanced Sampling=true, K (Number of samples)=102025.02 | 76.14 | |
| Mistral-7B-Instruct-v0.3Noise-Enhanced Sampling=false, K (Number of samples)=102025.02 | 75.85 | |
| NCIModel=Qwen-2.5-7B-Instruct, Single Sample=true2026.02 | 75.83 | |
| NCIBackbone=Qwen-3-32B2026.02 | 75.75 | |
| fDBDModel=Llama-3.2-3B-Instruct, Single Sample=true, k=ALL2026.02 | 75.59 | |
| LSModel=Qwen2.5-Math-1.5B-Instruct2025.08 | 75.57 | |
| CoE-CModel=Qwen-2.5-7B-Instruct, Single Sample=true2026.02 | 75.5 | |
| AttnScoreLLM=Llama3.1-8B2026.04 | 75.2 | |
| CoE-RModel=Qwen-2.5-7B-Instruct, Single Sample=true2026.02 | 75.13 | |
| Predictive ProbabilityBackbone=Qwen-3-32B2026.02 | 74.99 | |
| Self-FVModel=LLaMA2-7B2026.04 | 74.9 | |
| ARS (CCS)Model=DeepSeek-R1-Distill-Llama-8B, Single Sampling=true, Supervision=false2026.01 | 74.72 | |
| LN Predictive ProbabilityBackbone=Qwen-3-32B2026.02 | 74.38 | |
| AttnScoreLLM=Llama3.2-3B2026.04 | 74.3 | |
| SelfCheckGPT NLIModel=Llama-3.2-3B-Instruct, Single Sample=false2026.02 | 74.29 | |
| AttnScoreLLM=Phi3.52026.04 | 74.1 | |
| Multiple Hypothesis Testing for Hallucination DetectionModel=Llama-3.2-3B-Instruct2025.08 | 74.1 | |
| Max PModel=Llama-3.2-3B-Instruct, Single Sample=true2026.02 | 73.9 | |
| Lexical SimilarityModel=Llama-3.2-3B-Instruct, Single Sample=false2026.02 | 73.66 | |
| Predictive ProbabilityModel=Qwen-2.5-7B-Instruct, Single Sample=true2026.02 | 73.29 | |
| LN Predictive ProbabilityModel=Qwen-2.5-7B-Instruct, Single Sample=true2026.02 | 73.01 | |
| RACEModel=Qwen3-8B, Single Sampling=false, Supervision=false2026.01 | 72.55 | |
| Semantic EntropyModel=Qwen3-8B, Single Sampling=false, Supervision=false2026.01 | 72.51 | |
| Phi-3-mini-4k-instruct (3.8B)Noise-Enhanced Sampling=true, K (Number of samples)=102025.02 | 72.51 | |
| Lexical SimilarityModel=Qwen-2.5-7B-Instruct, Single Sample=false2026.02 | 72.02 | |
| FESSignal=Free-energy + SFF2026.06 | 72 | |
| AveragingModel=Llama-3.2-3B-Instruct2025.08 | 71.92 | |
| Function VectorModel=LLaMA2-7B2026.04 | 71.9 | |
| Llama-2-7B-chatNoise-Enhanced Sampling=false, K (Number of samples)=102025.02 | 71.56 | |
| PerplexityModel=Qwen-2.5-7B-Instruct, Single Sample=true2026.02 | 71.54 | |
| AttnScoreLLM=Mistral-Nemo2026.04 | 71.2 | |
| Predictive ProbabilityModel=Llama-3.2-3B-Instruct, Single Sample=true2026.02 | 70.88 | |
| LN Predictive ProbabilityModel=Llama-3.2-3B-Instruct, Single Sample=true2026.02 | 70.68 | |
| G-DetectorModel=DeepSeek-R1-Distill-Llama-8B, Single Sampling=true, Supervision=true2026.01 | 70.38 | |
| P(True)Model=Qwen-2.5-7B-Instruct, Single Sample=true2026.02 | 70.31 | |
| SelfCheckGPTModel=Meta-Llama-3-8B-Instruct, Category=Multi-trajectory / LLM-based2026.05 | 70.2 | |
| Best spec.Description=Strongest non-FES in-tree attention-spectral baseline2026.06 | 70 | |
| PerplexityModel=Llama-3.2-3B-Instruct, Single Sample=true2026.02 | 69.63 | |
| RACEModel=DeepSeek-R1-Distill-Llama-8B, Single Sampling=false, Supervision=false2026.01 | 68.59 | |
| Semantic EntropyModel=Qwen-2.5-7B-Instruct, Single Sample=false2026.02 | 66.83 | |
| Lexical SimilarityModel=Qwen3-8B, Single Sampling=false, Supervision=false2026.01 | 66.38 | |
| Phi-3-mini-4k-instruct (3.8B)Noise-Enhanced Sampling=false, K (Number of samples)=102025.02 | 65.86 | |
| UQ_ICLModel=LLaMA2-7B2026.04 | 65.7 | |
| AveragingModel=Qwen2.5-Math-1.5B-Instruct2025.08 | 64.7 | |
| Semantic EntropyModel=Llama-3.2-3B-Instruct, Single Sample=false2026.02 | 64.4 | |
| Self-FVModel=LLaMA2-13B2026.04 | 63.8 | |
| clustered_SEModel=Llama-3.2-3B-Instruct2025.08 | 63.61 | |
| Function VectorModel=LLaMA2-13B2026.04 | 62.7 | |
| Semantic EntropyModel=LLaMA2-7B2026.04 | 62.6 | |
| α_clustered_SEModel=Llama-3.2-3B-Instruct2025.08 | 62.34 | |
| Multiple Hypothesis Testing for Hallucination DetectionModel=Qwen2.5-Math-1.5B-Instruct2025.08 | 62.24 | |
| Semantic EntropyModel=DeepSeek-R1-Distill-Llama-8B, Single Sampling=false, Supervision=false2026.01 | 61.98 | |
| RHDModel=DeepSeek-R1-Distill-Llama-8B, Single Sampling=true, Supervision=false2026.01 | 61.67 | |
| UQ_ICLModel=LLaMA2-13B2026.04 | 61.6 | |
| clustered_SEModel=Qwen2.5-Math-1.5B-Instruct2025.08 | 61.24 | |
| PerplexityModel=Qwen3-8B, Single Sampling=true, Supervision=false2026.01 | 60.8 |