Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 44 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
1 benchmarks · 3 papers
Tool Use and Reasoning
1 benchmarks · 1 papers
Task performance (Cart Score)
1 benchmarks · 1 papers
Workflow Execution (Average)
1 benchmarks · 1 papers
Agentic Workflow Performance (Iterative Refinement Loops)
1 benchmarks · 1 papers
Agentic Workflow Performance (Retry Loops)
1 benchmarks · 1 papers
Agent Memory Management
1 benchmarks · 1 papers
Desktop UI Navigation
1 benchmarks · 1 papers
Agentic Workflow Performance (Static)
1 benchmarks · 1 papers
Sequential Task Solving
1 benchmarks · 1 papers
Agentic Reliability
1 benchmarks · 1 papers
Multimodal Web Navigation
1 benchmarks · 2 papers
Operating System Control
1 benchmarks · 1 papers
Tool-Need Prediction
1 benchmarks · 1 papers
Tool-Risk Prediction
1 benchmarks · 1 papers
Tool-call Monitoring
1 benchmarks · 1 papers
Agent-based Data Analysis
1 benchmarks · 1 papers
Tool-use runtime evaluation
1 benchmarks · 1 papers
Autonomous Incident Management
1 benchmarks · 1 papers
Terminal task resolution
1 benchmarks · 1 papers
Downstream skill execution
1 benchmarks · 1 papers
Web Browsing Action Prediction
1 benchmarks · 1 papers
Single-Agent Household Planning
1 benchmarks · 1 papers
Multi-turn planning
1 benchmarks · 1 papers
Single-Agent Web Navigation
1 benchmarks · 1 papers
Query routing and tool-calling accuracy evaluation
1 benchmarks · 1 papers
Agentic Workflow Injection Detection
1 benchmarks · 1 papers
Multi-agent game strategic reasoning
1 benchmarks · 1 papers
Strategic reasoning in auction games
1 benchmarks · 1 papers
Plan-level Tool Use
1 benchmarks · 1 papers
Multi-constraint search problem solving
1 benchmarks · 1 papers
Household Task Execution
1 benchmarks · 1 papers
Science Experiment Execution
Page 44 of 49
Previous
Next