Hallucination Detection on MATH 500
88.89AUROCEigenWD M1
Evaluation Results
| Method | Links | |
|---|---|---|
| EigenWD M1Target Model=DeepSeek-Chat, Teacher Model=Llama-2-7B, Detection Setting=Black-box case study2026.03 | 88.89 | |
| ARS (CCS)Model=DeepSeek-R1-Distill-Llama-8B, Judge Method=ROUGE2026.01 | 88 | |
| ERTarget Model=DeepSeek-Chat, Teacher Model=Llama-2-7B, Detection Setting=Black-box case study2026.03 | 87.6 | |
| ARS (CCS)Model=DeepSeek-R1-Distill-Llama-8B, Single Sampling=true, Supervision=false2026.01 | 86.38 | |
| ARS (CCS)Model=DeepSeek-R1-Distill-Llama-8B, Judge Method=DeepSeek-R1-Distill-Qwen-32B2026.01 | 86.24 | |
| TSVModel=DeepSeek-R1-Distill-Llama-8B, Judge Method=ROUGE2026.01 | 84.71 | |
| LSTarget Model=DeepSeek-Chat, Teacher Model=Llama-2-7B, Detection Setting=Black-box case study2026.03 | 83.35 | |
| TSVModel=DeepSeek-R1-Distill-Llama-8B, Judge Method=BLEURT2026.01 | 81.36 | |
| ARS (CCS)Model=DeepSeek-R1-Distill-Llama-8B, Judge Method=BLEURT2026.01 | 81 | |
| ARS (Probing)Model=DeepSeek-R1-Distill-Llama-8B, Single Sampling=true, Supervision=true2026.01 | 79.95 | |
| ARS (CCS)Model=Qwen3-8B, Single Sampling=true, Supervision=false2026.01 | 78.66 | |
| TSVModel=DeepSeek-R1-Distill-Llama-8B, Judge Method=DeepSeek-R1-Distill-Qwen-32B2026.01 | 78.64 | |
| ARS (Probing)Model=Qwen3-8B, Single Sampling=true, Supervision=true2026.01 | 78.17 | |
| G-DetectorModel=DeepSeek-R1-Distill-Llama-8B, Judge Method=ROUGE2026.01 | 71.48 | |
| FESSignal=Free-energy + SFF2026.06 | 68.3 | |
| Best spec.Description=Strongest non-FES in-tree attention-spectral baseline2026.06 | 67.4 | |
| G-DetectorModel=DeepSeek-R1-Distill-Llama-8B, Single Sampling=true, Supervision=true2026.01 | 64.45 | |
| TSVModel=DeepSeek-R1-Distill-Llama-8B, Single Sampling=true, Supervision=true2026.01 | 63.24 | |
| TSVModel=Qwen3-8B, Single Sampling=true, Supervision=true2026.01 | 63.12 | |
| RACEModel=Qwen3-8B, Single Sampling=false, Supervision=false2026.01 | 63.02 | |
| G-DetectorModel=DeepSeek-R1-Distill-Llama-8B, Judge Method=BLEURT2026.01 | 61.44 | |
| DSETarget Model=DeepSeek-Chat, Teacher Model=Llama-2-7B, Detection Setting=Black-box case study2026.03 | 59.97 | |
| Best LMDescription=Best language-model probability baseline among MSP and PPL−12026.06 | 59.6 | |
| SelfCKGPTModel=DeepSeek-R1-Distill-Llama-8B, Single Sampling=false, Supervision=false2026.01 | 59.15 | |
| Best rep.Description=Strongest reported reference hidden-state or sampling baseline2026.06 | 58 | |
| G-DetectorModel=Qwen3-8B, Single Sampling=true, Supervision=true2026.01 | 57.67 | |
| RHDModel=DeepSeek-R1-Distill-Llama-8B, Single Sampling=true, Supervision=false2026.01 | 56.5 | |
| Semantic EntropyModel=Qwen3-8B, Single Sampling=false, Supervision=false2026.01 | 56.13 | |
| ESTarget Model=DeepSeek-Chat, Teacher Model=Llama-2-7B, Detection Setting=Black-box case study2026.03 | 55.87 | |
| SelfCKGPTModel=Qwen3-8B, Single Sampling=false, Supervision=false2026.01 | 55.47 | |
| RACEModel=DeepSeek-R1-Distill-Llama-8B, Single Sampling=false, Supervision=false2026.01 | 53.55 | |
| LNETarget Model=DeepSeek-Chat, Teacher Model=Llama-2-7B, Detection Setting=Black-box case study2026.03 | 53.44 | |
| PerplexityModel=Qwen3-8B, Single Sampling=true, Supervision=false2026.01 | 51.62 | |
| RHDModel=Qwen3-8B, Single Sampling=true, Supervision=false2026.01 | 50.51 | |
| Lexical SimilarityModel=DeepSeek-R1-Distill-Llama-8B, Single Sampling=false, Supervision=false2026.01 | 49.92 | |
| Verbalized CertaintyModel=DeepSeek-R1-Distill-Llama-8B, Single Sampling=true, Supervision=false2026.01 | 48.98 | |
| Lexical SimilarityModel=Qwen3-8B, Single Sampling=false, Supervision=false2026.01 | 44.13 | |
| Semantic EntropyModel=DeepSeek-R1-Distill-Llama-8B, Single Sampling=false, Supervision=false2026.01 | 43.6 | |
| PerplexityModel=DeepSeek-R1-Distill-Llama-8B, Single Sampling=true, Supervision=false2026.01 | 40.96 | |
| G-DetectorModel=DeepSeek-R1-Distill-Llama-8B, Judge Method=DeepSeek-R1-Distill-Qwen-32B2026.01 | 38.06 | |
| Verbalized CertaintyModel=Qwen3-8B, Single Sampling=true, Supervision=false2026.01 | 23.87 |