LLM Judgement Confidence Estimation on TL;DR (test)
0.4269RKVerbalized Confidence
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Verbalized ConfidenceJudge Model=Mistral-7B2026.05 | 0.4269 | 0.5798 | |
| Predictive ProbabilityJudge Model=Mistral-7B2026.05 | 0.4191 | 0.5806 | |
| Predictive ProbabilityJudge Model=Llama3-70B2026.05 | 0.4159 | 0.5852 | |
| Random AnnotatorJudge Model=Mistral-7B2026.05 | 0.4085 | 0.601 | |
| Simulated AnnotatorsJudge Model=Mistral-7B2026.05 | 0.4029 | 0.5987 | |
| Learning Confidence (Vanilla)Judge Model=Mistral-7B2026.05 | 0.397 | 0.6081 | |
| Simulated AnnotatorsJudge Model=Llama3-70B2026.05 | 0.3919 | 0.6052 | |
| Random AnnotatorJudge Model=Llama3-70B2026.05 | 0.3895 | 0.6123 | |
| Predictive ProbabilityJudge Model=Qwen2.5-72B2026.05 | 0.3718 | 0.6271 | |
| Random AnnotatorJudge Model=Qwen2.5-72B2026.05 | 0.3646 | 0.6364 | |
| Simulated AnnotatorsJudge Model=Qwen2.5-72B2026.05 | 0.3613 | 0.631 | |
| Learning Confidence (Vanilla)Judge Model=Llama3-70B2026.05 | 0.3584 | 0.6401 | |
| Margin-Adaptive Confidence RankingJudge Model=Mistral-7B2026.05 | 0.3572 | 0.6409 | |
| Learning Confidence (Vanilla)Judge Model=Qwen2.5-72B2026.05 | 0.3382 | 0.6577 | |
| Margin-Adaptive Confidence RankingJudge Model=Llama3-70B2026.05 | 0.3137 | 0.6813 | |
| Margin-Adaptive Confidence RankingJudge Model=Qwen2.5-72B2026.05 | 0.279 | 0.7021 |