ResearchDatasetsHumanEval, MBPP, and BigCodeBenchFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsCode GenerationHumanEval+, MBPP+, and BigCodeBench Aggregate70.72Average Score12