LLM Judgement Confidence Estimation on AlpacaEval (test)
0.4367Rank Correlation (RK)Verbalized Confidence
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Verbalized ConfidenceJudge Model=Mistral-7B2026.05 | 0.4367 | 56.25 | |
| Random AnnotatorJudge Model=Mistral-7B2026.05 | 0.4321 | 56.71 | |
| Predictive ProbabilityJudge Model=Mistral-7B2026.05 | 0.4214 | 57.95 | |
| Simulated AnnotatorsJudge Model=Mistral-7B2026.05 | 0.4177 | 58.16 | |
| Predictive ProbabilityJudge Model=Qwen2.5-72B2026.05 | 0.4025 | 59.85 | |
| Predictive ProbabilityJudge Model=Llama3-70B2026.05 | 0.4022 | 59.87 | |
| Simulated AnnotatorsJudge Model=Qwen2.5-72B2026.05 | 0.3922 | 60.96 | |
| Random AnnotatorJudge Model=Qwen2.5-72B2026.05 | 0.392 | 60.84 | |
| Random AnnotatorJudge Model=Llama3-70B2026.05 | 0.3867 | 61.44 | |
| Learning Confidence (Vanilla)Judge Model=Mistral-7B2026.05 | 0.3865 | 61.79 | |
| Simulated AnnotatorsJudge Model=Llama3-70B2026.05 | 0.3839 | 61.62 | |
| Margin-Adaptive Confidence RankingJudge Model=Mistral-7B2026.05 | 0.3393 | 66.72 | |
| Learning Confidence (Vanilla)Judge Model=Qwen2.5-72B2026.05 | 0.337 | 65.89 | |
| Learning Confidence (Vanilla)Judge Model=Llama3-70B2026.05 | 0.3236 | 67.3 | |
| Margin-Adaptive Confidence RankingJudge Model=Llama3-70B2026.05 | 0.2776 | 70.48 | |
| Margin-Adaptive Confidence RankingJudge Model=Qwen2.5-72B2026.05 | 0.2707 | 70.64 |