ResearchDatasets10-domain aggregateFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsDialogic Deference Evaluation10-domain aggregate (including TruthfulQA, GPQA, HARP, SocialIQA, AdvisorQA, r/AIO, etc.) N=3,244 (test)69.6Average Accuracy C14
Dialogic Deference Evaluation10-domain aggregate (including TruthfulQA, GPQA, HARP, SocialIQA, AdvisorQA, r/AIO, etc.) N=3,244 (test)69.6Average Accuracy C14