ResearchDatasetsGSM8K, MATH500, MBPP+, HumanEval+FollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsMathematical Reasoning and Code GenerationGSM8K, MATH500, MBPP+, HumanEval+ Average81.03Accuracy12Inference EfficiencyGSM8K, MATH500, MBPP+, HumanEval+ Average9.34Avg. TPF4