ResearchTasksPredicting human judges' overall quality (Q0)FollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedSynthetic Conversations (5-fold cross-evaluation)Oracle0.611Pearson's r12Feb 26, 2026Real Human-Agent Conversations (test)Oracle0.717Pearson's rho10Feb 26, 2026