Hallucination detection on MMLU-Pro
87.08AUROCDRIFT
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| DRIFTModel=Qwen-2.5-7B-Instruct, Input=answer2026.01 | 87.08 | — | — | |
| DRIFTModel=Qwen-2.5-7B-Instruct, Input=question2026.01 | 82.21 | — | — | |
| HaloScopeModel=Qwen-2.5-7B-Instruct, Input=question2026.01 | 81.08 | — | — | |
| HaloScopeModel=LLaMA 2 Chat 7B, Input=question2026.01 | 79.56 | — | — | |
| DRIFTModel=LLaMA 2 Chat 7B, Input=answer2026.01 | 77.85 | — | — | |
| HaloScopeModel=Gemma-3-4b-it, Input=question2026.01 | 77.73 | — | — | |
| HaloScopeModel=Qwen-2.5-7B-Instruct, Input=answer2026.01 | 77.56 | — | — | |
| HaloScopeModel=Gemma-3-4b-it, Input=answer2026.01 | 77.56 | — | — | |
| Semantic EntropyModel=LLaMA 2 Chat 7B2026.01 | 76.76 | — | — | |
| DRIFTModel=Gemma-3-4b-it, Input=answer2026.01 | 75.85 | — | — | |
| Attention probeModel=Qwen2.5-7B, Generation Stage=Post-generation2026.06 | 75.42 | — | — | |
| HaloScopeModel=LLaMA 2 Chat 7B, Input=answer2026.01 | 74.83 | — | — | |
| DRIFTModel=Gemma-3-4b-it, Input=question2026.01 | 74.65 | — | — | |
| Attention probeModel=Qwen2.5-7B, Generation Stage=Pre-generation2026.06 | 74.4 | — | — | |
| Semantic EntropyModel=Qwen-2.5-7B-Instruct2026.01 | 74.05 | — | — | |
| EntropyModel=Qwen2.5-7B, Generation Stage=Post-generation2026.06 | 72.32 | — | — | |
| DRIFTModel=LLaMA 2 Chat 7B, Input=question2026.01 | 71.22 | — | — | |
| Linear probeModel=Qwen2.5-7B, Generation Stage=Pre-generation2026.06 | 69.23 | — | — | |
| Linear probeModel=Qwen2.5-7B, Generation Stage=Post-generation2026.06 | 69.18 | — | — | |
| EntropyModel=Qwen2.5-3B, Generation Stage=Post-generation2026.06 | 64.99 | — | — | |
| Linear probeModel=Qwen2.5-3B, Generation Stage=Post-generation2026.06 | 64.03 | — | — | |
| Linear probeModel=Qwen2.5-3B, Generation Stage=Pre-generation2026.06 | 63.89 | — | — | |
| Self-check, zero-shotModel=Qwen2.5-7B, Generation Stage=Pre-generation2026.06 | 60.2 | — | — | |
| Question lengthModel=Qwen2.5-7B, Generation Stage=Pre-generation2026.06 | 60.15 | — | — | |
| Question lengthModel=Qwen2.5-3B, Generation Stage=Pre-generation2026.06 | 59.21 | — | — | |
| Semantic EntropyModel=Gemma-3-4b-it2026.01 | 57.92 | — | — | |
| Attention probe, soft targetModel=Qwen2.5-7B, Generation Stage=Pre-generation2026.06 | 56.95 | — | — | |
| Self-check, zero-shotModel=Qwen2.5-3B, Generation Stage=Pre-generation2026.06 | 53.73 | — | — | |
| Attention probe, soft targetModel=Qwen2.5-3B, Generation Stage=Pre-generation2026.06 | 53.55 | — | — | |
| Attention probeModel=Qwen2.5-3B, Generation Stage=Post-generation2026.06 | 50 | — | — | |
| Attention probeModel=Qwen2.5-3B, Generation Stage=Pre-generation2026.06 | 50 | — | — | |
| FACTOOLModels=Qwen2.5-7B2025.11 | — | 59.6 | 69.23 | |
| FACTOOLModels=Llama3.1-8B2025.11 | — | 55.33 | 68.84 | |
| FACTOOLModels=Mistral-8B2025.11 | — | 62.26 | 71.43 | |
| FIREModels=Qwen2.5-7B2025.11 | — | 61.05 | 69.23 | |
| FIREModels=Llama3.1-8B2025.11 | — | 58.85 | 69.51 | |
| FIREModels=Mistral-8B2025.11 | — | 58.23 | 71.11 | |
| HaluAgentModels=Qwen2.5-7B2025.11 | — | 54.68 | 55 | |
| HaluAgentModels=Llama3.1-8B2025.11 | — | 54.41 | 54.05 | |
| HaluAgentModels=Mistral-8B2025.11 | — | 54.29 | 70.37 | |
| LEAPModels=Qwen2.5-7B2025.11 | — | 69.81 | 75.31 | |
| LEAPModels=Llama3.1-8B2025.11 | — | 64.23 | 71.18 | |
| LEAPModels=Mistral-8B2025.11 | — | 63.21 | 71.15 | |
| LN-entropyModels=Qwen2.5-7B2025.11 | — | 56.67 | 67.5 | |
| LN-entropyModels=Llama3.1-8B2025.11 | — | 58.34 | 68.51 | |
| LN-entropyModels=Mistral-8B2025.11 | — | 58.67 | 68.04 | |
| PerplexityModels=Qwen2.5-7B2025.11 | — | 57.33 | 68 | |
| PerplexityModels=Llama3.1-8B2025.11 | — | 55 | 70.97 | |
| PerplexityModels=Mistral-8B2025.11 | — | 56.67 | 67.17 | |
| SAFEModels=Qwen2.5-7B2025.11 | — | 55.81 | 68.28 | |
| SAFEModels=Llama3.1-8B2025.11 | — | 57.96 | 70.49 | |
| SAFEModels=Mistral-8B2025.11 | — | 60.61 | 72.19 | |
| Self-Check(0)Models=Qwen2.5-7B2025.11 | — | 56.66 | 69.54 | |
| Self-Check(0)Models=Llama3.1-8B2025.11 | — | 54.88 | 70.74 | |
| Self-Check(0)Models=Mistral-8B2025.11 | — | 60.14 | 70.37 | |
| Self-Check(3)Models=Qwen2.5-7B2025.11 | — | 59.56 | 64.52 | |
| Self-Check(3)Models=Llama3.1-8B2025.11 | — | 55.25 | 70.67 | |
| Self-Check(3)Models=Mistral-8B2025.11 | — | 59.12 | 70.42 | |
| Semantic EntropyModels=Qwen2.5-7B2025.11 | — | 55.67 | 66.67 | |
| Semantic EntropyModels=Llama3.1-8B2025.11 | — | 55.33 | 65.59 | |
| Semantic EntropyModels=Mistral-8B2025.11 | — | 54.67 | 69.64 |