ResearchTasksLarge Language Model Reasoning and CodingFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedGPQA-Diamond, HumanEval, LiveCodeBench v6, AIME25, and MATH500BDR (Ours)88.3Throughput (Tokens/User)10Apr 23, 2026
GPQA-Diamond, HumanEval, LiveCodeBench v6, AIME25, and MATH500BDR (Ours)88.3Throughput (Tokens/User)10Apr 23, 2026