LLM Judgement Confidence Estimation on HH-RLHF (test)
0.4763RKPredictive Probability
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Predictive ProbabilityJudge Model=Mistral-7B2026.05 | 0.4763 | 0.5259 | |
| Verbalized ConfidenceJudge Model=Mistral-7B2026.05 | 0.4738 | 0.523 | |
| Predictive ProbabilityJudge Model=Llama3-70B2026.05 | 0.4478 | 0.551 | |
| Predictive ProbabilityJudge Model=Qwen2.5-72B2026.05 | 0.4409 | 0.554 | |
| Simulated AnnotatorsJudge Model=Mistral-7B2026.05 | 0.3986 | 0.5999 | |
| Learning Confidence (Vanilla)Judge Model=Mistral-7B2026.05 | 0.3841 | 0.6154 | |
| Random AnnotatorJudge Model=Mistral-7B2026.05 | 0.3803 | 0.6202 | |
| Random AnnotatorJudge Model=Qwen2.5-72B2026.05 | 0.376 | 0.6369 | |
| Random AnnotatorJudge Model=Llama3-70B2026.05 | 0.3751 | 0.6374 | |
| Simulated AnnotatorsJudge Model=Qwen2.5-72B2026.05 | 0.3678 | 0.6473 | |
| Learning Confidence (Vanilla)Judge Model=Llama3-70B2026.05 | 0.3597 | 0.6489 | |
| Simulated AnnotatorsJudge Model=Llama3-70B2026.05 | 0.3571 | 0.6537 | |
| Learning Confidence (Vanilla)Judge Model=Qwen2.5-72B2026.05 | 0.3482 | 0.6578 | |
| Margin-Adaptive Confidence RankingJudge Model=Mistral-7B2026.05 | 0.3286 | 0.6805 | |
| Margin-Adaptive Confidence RankingJudge Model=Llama3-70B2026.05 | 0.3094 | 0.6945 | |
| Margin-Adaptive Confidence RankingJudge Model=Qwen2.5-72B2026.05 | 0.2813 | 0.7126 |