ResearchDatasetsSpeechJudgeFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsPreference EvaluationSpeechJudge75Acc@0.515Pairwise preference predictionSpeechJudge human preference (test)70.4Accuracy4