ResearchTasksReward Model LearningFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedUFBURM72.5Win Rate10Jul 7, 2026