ResearchTasksMathematical Reasoning and Code GenerationFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedGSM8K, MATH500, MBPP+, HumanEval+ AverageMBD-LLaDA2-Mini81.03Accuracy12Jun 30, 2026Aggregate (MATH-500, AIME-2024, GSM8K, HumanEval+, MBPP+)DeepSeek-R1-Distill-Qwen-32B83.17Average Score10Mar 27, 2026Mean (GSM8K, MATH, HumanEval, MBPP)FeF-DLLM52.06Accuracy7May 15, 2026
Aggregate (MATH-500, AIME-2024, GSM8K, HumanEval+, MBPP+)DeepSeek-R1-Distill-Qwen-32B83.17Average Score10Mar 27, 2026