ResearchTasksPairwise EvaluationFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedBIGGENHuman-crafted (existing) Rubrics78.33Human Agreement41Jun 1, 2026AlpacaEvalFine-tuned Rubric Generator72.4Human Agreement37Jun 1, 2026MT-BenchFine-tuned Rubric Generator83.69Human Agreement Rate9Jun 1, 2026HH-RLHF (test)pairwise evaluator95.2Test Accuracy4Apr 14, 2026