ResearchTasksReasoning Episode ClassificationFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedOmni-MATH human-annotated Reasoning episodes (gold set)GPT-586.33Accuracy8Feb 26, 2026Omni-MATH Non-Reasoning episodes (human-annotated gold set)GPT-4.189.34Accuracy4Feb 26, 2026