Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 43 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
1 benchmarks · 1 papers
Web-search agent tasks
1 benchmarks · 1 papers
Safe Agent Planning
1 benchmarks · 1 papers
Agent Tool-use and Reasoning
1 benchmarks · 1 papers
GUI State Control
1 benchmarks · 1 papers
Code-interpreter tasks
1 benchmarks · 1 papers
World Model Task Execution
1 benchmarks · 1 papers
Adversarial Attack against WebExperT agent
1 benchmarks · 1 papers
Planning and self-evolution under curriculum setting
1 benchmarks · 1 papers
Earth Observation Agent Task Completion
1 benchmarks · 1 papers
LLM Agent Safety
1 benchmarks · 1 papers
Tool-use performance
1 benchmarks · 1 papers
Text-based embodied task completion
1 benchmarks · 3 papers
Tool Use and Reasoning
1 benchmarks · 1 papers
Task performance (Cart Score)
1 benchmarks · 1 papers
Workflow Execution (Average)
1 benchmarks · 1 papers
Agentic Workflow Performance (Iterative Refinement Loops)
1 benchmarks · 1 papers
Agentic Workflow Performance (Retry Loops)
1 benchmarks · 1 papers
Response and Tool-Call Quality Evaluation
1 benchmarks · 1 papers
Proactive Agent Task Execution
1 benchmarks · 1 papers
Tool-use safety validation
1 benchmarks · 1 papers
Web GUI Interaction
1 benchmarks · 1 papers
Agentic Workflow Performance (Static)
1 benchmarks · 1 papers
Sequential Task Solving
1 benchmarks · 2 papers
Operating System Control
1 benchmarks · 1 papers
Tool-Need Prediction
1 benchmarks · 1 papers
Tool-Risk Prediction
1 benchmarks · 1 papers
Tool-call Monitoring
1 benchmarks · 1 papers
Tool-use runtime evaluation
1 benchmarks · 1 papers
Autonomous Incident Management
1 benchmarks · 1 papers
Terminal task resolution
1 benchmarks · 1 papers
Downstream skill execution
1 benchmarks · 1 papers
Web Browsing Action Prediction
Page 43 of 49
Previous
Next