ResearchDatasetsReasoning Agreement BenchmarkFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsLLM Judging AgreementReasoning Agreement Benchmark 500-sample human-annotated0.91Cohen's Kappa15LLM Judging AgreementReasoning Agreement Benchmark 2,500 samples100Parsing Success Rate15