ResearchDatasetsHuggingGPT Human Evaluation SetFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsTask PlanningHuggingGPT Human Evaluation Set 130 diverse requests (test)0.912Passing Rate3Model SelectionHuggingGPT Human Evaluation Set 130 diverse requests (test)93.89Passing Rate1