Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 14 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
2 benchmarks · 4 papers
Agent Capability Evaluation
2 benchmarks · 2 papers
Autonomous Software Engineering
2 benchmarks · 1 papers
Agent Planning Security and Autonomy
2 benchmarks · 1 papers
Function Search
2 benchmarks · 1 papers
Item Crafting and Interaction Tasks
2 benchmarks · 1 papers
Tidy Up Room
2 benchmarks · 2 papers
Agentic Terminal Tasks
2 benchmarks · 1 papers
Human-Agent Coordination
2 benchmarks · 1 papers
Multi-agent Task Fulfillment
2 benchmarks · 1 papers
Long-Horizon Agent Tasks
2 benchmarks · 2 papers
Multi-agent Simulation
2 benchmarks · 1 papers
Throughput Efficiency
2 benchmarks · 1 papers
Agentic Output Verification
2 benchmarks · 1 papers
Inquire-driven GUI Navigation
2 benchmarks · 1 papers
GUI Execution
2 benchmarks · 1 papers
Collaborative software engineering
2 benchmarks · 2 papers
Malicious Agent
2 benchmarks · 1 papers
Skill Evaluation
2 benchmarks · 1 papers
Tool Completion
2 benchmarks · 2 papers
Multi-tool calling
2 benchmarks · 1 papers
Browser Automation
2 benchmarks · 1 papers
Sub-task Completion
2 benchmarks · 1 papers
Tool Sequence Recommendation
2 benchmarks · 1 papers
Multi-Agent Construction
2 benchmarks · 1 papers
Web-based Agent Task Completion
2 benchmarks · 1 papers
Sub-workflow Execution
2 benchmarks · 1 papers
Personalized Agentic Social Support
2 benchmarks · 1 papers
Single-turn Tool Calling
2 benchmarks · 1 papers
Slide Automation
2 benchmarks · 1 papers
Multi-agent Semantic Object Navigation
2 benchmarks · 1 papers
Website Navigation
2 benchmarks · 1 papers
Safety and Action Blocking
Page 14 of 49
Previous
Next