ResearchTasksAgent Capability EvaluationFollowBenchmarksDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyDataset NameSOTA methodMetricTrendResultsLast UpdatedSEAL 0MiroThinker-H161.3Average Score (@8)19Apr 23, 2026ACEBench AgentGPT-4.195Multi-Step Reasoning Score13Feb 26, 2026