ResearchDatasetsReasoning BenchmarksFollowBenchmarksTask NameDataset NameSOTA MethodSortMost resultsRecently updatedMost papersApplyTask NameDataset NameSOTA ResultTrendResultsMulti-agent ReasoningReasoning Benchmarks Cooperative AutoGen framework (test)83.58Overall Accuracy2Multi-agent ReasoningReasoning Benchmarks Competitive MAD framework (test)0.851Average Score2Page 2 of 2PreviousNext