ResearchDatasetsResponse Harmfulness Detection BenchmarksFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsResponse Harmfulness DetectionResponse Harmfulness Detection Benchmarks (HarmBench, SafeRLHF, BeaverTails, XSTest, WildGuard)0.833Macro Avg F121
Response Harmfulness DetectionResponse Harmfulness Detection Benchmarks (HarmBench, SafeRLHF, BeaverTails, XSTest, WildGuard)0.833Macro Avg F121