Uncertainty Estimation on TruthfulQA
26.8PRRInternal Confidence
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Internal ConfidenceModel=Llama-8B2025.06 | 26.8 | 63.2 | 15.4 | |
| Min-K EntropyModel=Qwen-14B2025.06 | 22.8 | 58.1 | — | |
| Predictive EntropyModel=Qwen-14B2025.06 | 21.1 | 59.9 | — | |
| Min-K EntropyModel=Llama-8B2025.06 | 21 | 57.8 | — | |
| Internal ConfidenceModel=Qwen-14B2025.06 | 16.4 | 58 | 0.5 | |
| Predictive EntropyModel=Phi-3.8B2025.06 | 15.6 | 55.7 | — | |
| Min-K EntropyModel=Phi-3.8B2025.06 | 13.9 | 55.1 | — | |
| Predictive EntropyModel=Llama-8B2025.06 | 13.7 | 60 | — | |
| Internal ConfidenceModel=Phi-3.8B2025.06 | 13.2 | 56.4 | 40.7 | |
| Attentional EntropyModel=Llama-8B2025.06 | 9.6 | 53.3 | — | |
| Attentional EntropyModel=Phi-3.8B2025.06 | 8.3 | 52.4 | — | |
| PerplexityModel=Qwen-14B2025.06 | 6 | 52.6 | — | |
| PerplexityModel=Llama-8B2025.06 | 4.8 | 54.3 | — | |
| Max(− log p)Model=Qwen-14B2025.06 | 4.7 | 51 | — | |
| Max(− log p)Model=Llama-8B2025.06 | 4.1 | 52.4 | — | |
| PerplexityModel=Phi-3.8B2025.06 | 3.8 | 54.1 | — | |
| Attentional EntropyModel=Qwen-14B2025.06 | 3.2 | 50.7 | — | |
| P(YES) (naive avg)Model=Llama-8B2025.06 | 1.9 | 47 | 3.4 | |
| GLU-EDLModel=Qwen2026.06 | 0.392 | — | — | |
| LogTokUModel=Qwen2026.06 | 0.377 | — | — | |
| GLU-AUModel=Fanar2026.06 | 0.347 | — | — | |
| LogProbModel=Fanar2026.06 | 0.31 | — | — | |
| GLUModel=Fanar2026.06 | 0.307 | — | — | |
| RAUQModel=Fanar2026.06 | 0.282 | — | — | |
| P(true)Model=Gemma2026.06 | 0.267 | — | — | |
| GLUModel=Qwen2026.06 | 0.195 | — | — | |
| LogTokUModel=Fanar2026.06 | 0.192 | — | — | |
| LogProbModel=Qwen2026.06 | 0.185 | — | — | |
| P(true)Model=Fanar2026.06 | 0.181 | — | — | |
| Add-SαModel=Gemma2026.06 | 0.154 | — | — | |
| LogTokUModel=Gemma2026.06 | 0.115 | — | — | |
| P(true)Model=Qwen2026.06 | 0.102 | — | — | |
| LogProbModel=Gemma2026.06 | 0.091 | — | — | |
| GLUModel=Gemma2026.06 | 0.087 | — | — | |
| RAUQModel=Qwen2026.06 | 0.063 | — | — | |
| RAUQModel=Gemma2026.06 | -0.094 | — | — | |
| Max(− log p)Model=Phi-3.8B2025.06 | -1.3 | 50.3 | — | |
| P(YES) (naive avg)Model=Phi-3.8B2025.06 | -2 | 49.3 | 25.5 | |
| P(YES) (naive avg)Model=Qwen-14B2025.06 | -3.5 | 49.7 | 8.6 | |
| P(YES) (top right)Model=Llama-8B2025.06 | -11.9 | 43.3 | 55.7 | |
| P(YES) (top right)Model=Phi-3.8B2025.06 | -16.2 | 39.6 | 27.1 | |
| P(YES) (top right)Model=Qwen-14B2025.06 | -19 | 42.6 | 53.4 | |
| AleatoricBackbone=Mistral-7B2026.04 | — | 52.8 | — | |
| Closeness CentralityBackbone=Mistral-7B2026.04 | — | 63.9 | — | |
| Kernel Lang. Ent.Backbone=Mistral-7B2026.04 | — | 63.7 | — | |
| Max Sequence Prob.Backbone=Mistral-7B2026.04 | — | 53.4 | — | |
| Max Token Prob.Backbone=Mistral-7B2026.04 | — | 52.8 | — | |
| Mean Token EntropyBackbone=Mistral-7B2026.04 | — | 54.4 | — | |
| PerplexityBackbone=Mistral-7B2026.04 | — | 52.7 | — | |
| PTrueBackbone=Mistral-7B2026.04 | — | 45.1 | — | |
| SC + VCBackbone=Mistral-7B2026.04 | — | 58.1 | — | |
| SC Based VCBackbone=Mistral-7B2026.04 | — | 60.2 | — | |
| SC ScoreBackbone=Mistral-7B2026.04 | — | 60.3 | — | |
| Self CertaintyBackbone=Mistral-7B2026.04 | — | 58.6 | — | |
| SelfCheckGPTBackbone=Mistral-7B2026.04 | — | 53.4 | — | |
| SemanticEntropyBackbone=Mistral-7B2026.04 | — | 51.4 | — | |
| Token EntropyBackbone=Mistral-7B2026.04 | — | 54.2 | — | |
| TotalBackbone=Mistral-7B2026.04 | — | 57.9 | — |