ResearchDatasetsAnthropic-SafeRLHFFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsPreference evaluationAnthropic-SafeRLHF (target)41.7Win Rate2Preference evaluationAnthropic-SafeRLHF benchmark33.7Win Rate2