ResearchDatasetsPreferenceBenchFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsLLM-as-a-JudgePreferenceBench90.71Accuracy59LLM-as-a-JudgePreferenceBench90.2Accuracy21LLM-as-a-Judge Evaluation ConsistencyPreferenceBench79.73Kappa4