Uncertainty Estimation on SimpleQA, MuSiQue, and TruthfulQA Average
61AUROCInternal Confidence
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Internal ConfidenceModel=Phi-3.8B2025.06 | 61 | 23.2 | 22.7 | |
| Internal ConfidenceModel=Llama-8B2025.06 | 61 | 22.7 | 52.4 | |
| Predictive EntropyModel=Phi-3.8B2025.06 | 58.4 | 17 | — | |
| P(YES) (naive avg)Model=Phi-3.8B2025.06 | 58.2 | 15.1 | 52.9 | |
| Internal ConfidenceModel=Qwen-14B2025.06 | 57.2 | 14.7 | 5.4 | |
| Min-K EntropyModel=Phi-3.8B2025.06 | 56.3 | 16.3 | — | |
| P(YES) (top right)Model=Phi-3.8B2025.06 | 55.4 | 10.5 | 18.5 | |
| Predictive EntropyModel=Llama-8B2025.06 | 55.2 | 8.6 | — | |
| P(YES) (naive avg)Model=Qwen-14B2025.06 | 55.2 | 10 | 6.4 | |
| P(YES) (naive avg)Model=Llama-8B2025.06 | 55 | 13.2 | 16.8 | |
| Predictive EntropyModel=Qwen-14B2025.06 | 54.7 | 10.6 | — | |
| Min-K EntropyModel=Qwen-14B2025.06 | 54.7 | 11.1 | — | |
| Min-K EntropyModel=Llama-8B2025.06 | 54.5 | 12.1 | — | |
| Attentional EntropyModel=Llama-8B2025.06 | 54.1 | 7.7 | — | |
| PerplexityModel=Phi-3.8B2025.06 | 53.9 | 5.6 | — | |
| P(YES) (top right)Model=Llama-8B2025.06 | 53.7 | 6.9 | 69.7 | |
| PerplexityModel=Llama-8B2025.06 | 52.9 | 2.3 | — | |
| Attentional EntropyModel=Phi-3.8B2025.06 | 52.4 | 8.3 | — | |
| PerplexityModel=Qwen-14B2025.06 | 52.1 | 4.8 | — | |
| P(YES) (top right)Model=Qwen-14B2025.06 | 52 | 2.3 | 37.1 | |
| Max(− log p)Model=Llama-8B2025.06 | 51.9 | 2.5 | — | |
| Max(− log p)Model=Qwen-14B2025.06 | 51.4 | 3.4 | — | |
| Attentional EntropyModel=Qwen-14B2025.06 | 51.4 | 4.5 | — | |
| Max(− log p)Model=Phi-3.8B2025.06 | 51.3 | 3 | — |