ResearchTasksMalicious conversation detectionFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedMulti-turn Conversations 10,654 (full evaluation)peak + accumulation scoring90.8Recall1Feb 26, 2026