Predicting human judges' overall quality (Q0) on Real Human-Agent Conversations (test)
0.717Pearson's rhoOracle
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| OracleDescription=Includes actual judge response to Qi2024.12 | 0.717 | 0.289 | 0.711 | 0.675 | |
| Oracle + Personalized isotonic regressionConfiguration=Personalized isotonic regression correction2024.12 | 0.65 | 0.302 | 0.653 | 0.644 | |
| Oracle w/o LLM probsAblation=Withholding LLM response vector2024.12 | 0.625 | 0.357 | 0.629 | 0.599 | |
| Oracle w/o Personalized CalibrationAblation=Depersonalizing calibration network2024.12 | 0.582 | 0.389 | 0.587 | 0.565 | |
| LLM-RUBRIC2024.12 | 0.35 | 0.422 | 0.347 | 0.331 | |
| FActScore2024.12 | 0.216 | — | 0.218 | 0.207 | |
| Calibrated LLM Q02024.12 | 0.211 | 0.784 | 0.218 | 0.192 | |
| Expected LLM Q02024.12 | 0.143 | 0.901 | 0.141 | 0.138 | |
| Argmax LLM Q02024.12 | 0.106 | 1.186 | 0.123 | 0.12 | |
| Random Eval2024.12 | 0.011 | 1.427 | 0.006 | 0.005 |