ResearchTasksPreference-based AlignmentFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedStanford SHP (test)DPO83.5Win Rate (Llama)9Apr 28, 2026