ResearchDatasetsFairJudgeFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsLLM-as-a-JudgeFairJudge Benchmark 1K (test)71.5Agreement13