ResearchTasksUnsafe Prompt DetectionFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedToxicChat (test)OpenAI Moderation API0.815Precision16Feb 26, 2026XSTest (test)OpenAI Moderation API87.8Precision7Feb 26, 2026XSTestGradSafe-Zero93.6AUPRC4Feb 26, 2026ToxicChatGradSafe-Zero75.5AUPRC4Feb 26, 2026