Predicting human judges' overall quality (Q0) on Synthetic Conversations (5-fold cross-evaluation)
0.611Pearson's rOracle
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| OracleDescription=Includes actual judge response to Qi2024.12 | 0.611 | 0.237 | 0.626 | 0.605 | |
| Oracle w/o LLM probsAblation=Withholding LLM response vector2024.12 | 0.551 | 0.276 | 0.548 | 0.533 | |
| Oracle + Personalized isotonic regressionConfiguration=Personalized isotonic regression correction2024.12 | 0.521 | 0.273 | 0.526 | 0.519 | |
| Depersonalized Oracle + Personalized isotonic regressionConfiguration=Personalized isotonic regression correction2024.12 | 0.482 | 0.321 | 0.485 | 0.477 | |
| Oracle w/o Personalized CalibrationAblation=Depersonalizing calibration network2024.12 | 0.476 | 0.401 | 0.471 | 0.468 | |
| LLM-RUBRIC2024.12 | 0.401 | 0.396 | 0.398 | 0.393 | |
| Depersonalized OracleDescription=Uses distribution of responses of all other judges2024.12 | 0.362 | 0.492 | 0.355 | 0.338 | |
| FActScore2024.12 | 0.204 | — | 0.211 | 0.2 | |
| Calibrated LLM Q02024.12 | 0.198 | 0.801 | 0.196 | 0.193 | |
| Expected LLM Q02024.12 | 0.182 | 0.856 | 0.217 | 0.168 | |
| Argmax LLM Q02024.12 | 0.153 | 0.984 | 0.161 | 0.147 | |
| Random Eval2024.12 | 0.002 | 1.499 | -0.003 | -0.003 |