LLM Judgement Confidence Estimation on Chatbot Arena (test)
0.3418RKVerbalized Confidence
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Verbalized ConfidenceJudge Model=Mistral-7B2026.05 | 0.3418 | 65.03 | |
| Random AnnotatorJudge Model=Mistral-7B2026.05 | 0.3355 | 66.28 | |
| Predictive ProbabilityJudge Model=Mistral-7B2026.05 | 0.3323 | 66.5 | |
| Simulated AnnotatorsJudge Model=Mistral-7B2026.05 | 0.323 | 67.67 | |
| Learning Confidence (Vanilla)Judge Model=Mistral-7B2026.05 | 0.2817 | 70.32 | |
| Margin-Adaptive Confidence RankingJudge Model=Mistral-7B2026.05 | 0.2743 | 71.27 | |
| Simulated AnnotatorsJudge Model=Llama3-70B2026.05 | 0.2646 | 73.54 | |
| Random AnnotatorJudge Model=Llama3-70B2026.05 | 0.2629 | 73.41 | |
| Predictive ProbabilityJudge Model=Llama3-70B2026.05 | 0.2552 | 74.56 | |
| Learning Confidence (Vanilla)Judge Model=Llama3-70B2026.05 | 0.2486 | 75.2 | |
| Random AnnotatorJudge Model=Qwen2.5-72B2026.05 | 0.248 | 75.19 | |
| Simulated AnnotatorsJudge Model=Qwen2.5-72B2026.05 | 0.2469 | 75.12 | |
| Predictive ProbabilityJudge Model=Qwen2.5-72B2026.05 | 0.2457 | 75.36 | |
| Learning Confidence (Vanilla)Judge Model=Qwen2.5-72B2026.05 | 0.2435 | 76.1 | |
| Margin-Adaptive Confidence RankingJudge Model=Llama3-70B2026.05 | 0.2165 | 78.72 | |
| Margin-Adaptive Confidence RankingJudge Model=Qwen2.5-72B2026.05 | 0.2077 | 78.4 |