Turn-level correlation with human ratings on MultiChallenge
0.74Spearman CorrelationSKG-Eval
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| SKG-Eval2026.05 | 0.74 | 0.56 | |
| GPT-4o-Judgemode=history-aware2026.05 | 0.66 | 0.49 | |
| GPT-4o-Judgemode=turn-only2026.05 | 0.57 | 0.42 | |
| DeepEvalmodel=GPT-4o2026.05 | 0.55 | 0.4 | |
| LLM-Eval2026.05 | 0.49 | 0.35 | |
| ECoh2026.05 | 0.43 | 0.31 |