ResearchTasksBilevel Reinforcement Learning from Human FeedbackFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedBilevel RLHF Frameworks TheoreticalSAIL-RevKL2Local Sample Complexity3Jul 1, 2026