Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
227 benchmarks · 111 papers
Generative Modeling
214 benchmarks · 152 papers
Mathematical Problem Solving
138 benchmarks · 103 papers
Tool Use
115 benchmarks · 88 papers
Function Calling
76 benchmarks · 75 papers
Web navigation
75 benchmarks · 37 papers
Tool Calling
71 benchmarks · 20 papers
Task Planning
54 benchmarks · 16 papers
Tool Retrieval
52 benchmarks · 34 papers
GUI Navigation
43 benchmarks · 21 papers
World Modeling
41 benchmarks · 14 papers
Agentic Search
31 benchmarks · 13 papers
Tool selection
30 benchmarks · 22 papers
Agentic Reasoning
29 benchmarks · 29 papers
Agentic Tool-use
28 benchmarks · 14 papers
Deep search
28 benchmarks · 18 papers
Agentic Coding
25 benchmarks · 7 papers
Zero-shot Coordination
24 benchmarks · 11 papers
Prompt Injection Attack
22 benchmarks · 30 papers
General AI Assistant Tasks
22 benchmarks · 10 papers
Agentic Task
21 benchmarks · 2 papers
human-robot task planning and allocation
20 benchmarks · 14 papers
Computer Use
18 benchmarks · 1 papers
Targeted Web Crawling
18 benchmarks · 12 papers
GUI Automation
17 benchmarks · 9 papers
Agent Task Completion
17 benchmarks · 2 papers
LTL Instruction Following
17 benchmarks · 7 papers
Task success
16 benchmarks · 8 papers
Agentic
16 benchmarks · 10 papers
Agentic Task Completion
16 benchmarks · 6 papers
Failure attribution
16 benchmarks · 12 papers
Web Agent Navigation
16 benchmarks · 18 papers
Agent Task
Page 1 of 49
Previous
Next