ResearchDatasetsReal Human-Agent ConversationsFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsPredicting human judges' overall quality (Q0)Real Human-Agent Conversations (test)0.717Pearson's rho10
Predicting human judges' overall quality (Q0)Real Human-Agent Conversations (test)0.717Pearson's rho10