ResearchDatasetsTool benchmarkFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsReasoning Quality Evaluation120-tool benchmark 500 tasks simulated4.43Mean Score5