Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 13 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
3 benchmarks · 1 papers
Tool / Agent
3 benchmarks · 1 papers
Web Navigation Task Completion
3 benchmarks · 1 papers
Agentic execution trace analysis
3 benchmarks · 1 papers
Multi-agent Planning and Composition
3 benchmarks · 4 papers
Trip Planning
3 benchmarks · 4 papers
Multi-hop tool use
3 benchmarks · 3 papers
Downstream evaluation
3 benchmarks · 1 papers
Step-level tool invocation safety detection
3 benchmarks · 1 papers
Automated agent rollback
3 benchmarks · 1 papers
Web-based task completion
3 benchmarks · 1 papers
Web-based Agent Reasoning
3 benchmarks · 1 papers
Monkey-banana task execution
3 benchmarks · 1 papers
Tool-use Factuality Evaluation
3 benchmarks · 1 papers
Repeated Negotiation
3 benchmarks · 6 papers
Tool-use Agent Performance
3 benchmarks · 1 papers
Agent Trajectory Performance
3 benchmarks · 1 papers
Tool Reasoning
3 benchmarks · 2 papers
GUI Agent Navigation and Action
3 benchmarks · 1 papers
Head-to-Head Evaluation
3 benchmarks · 1 papers
Continual-memory deployment
3 benchmarks · 1 papers
Orchestration
3 benchmarks · 1 papers
Task Planning and Program Generation
3 benchmarks · 1 papers
Next-Action
3 benchmarks · 1 papers
Desktop GUI Task Success
3 benchmarks · 4 papers
Tool
2 benchmarks · 1 papers
Inquire-driven GUI Navigation
2 benchmarks · 1 papers
GUI Execution
2 benchmarks · 1 papers
Collaborative software engineering
2 benchmarks · 1 papers
Tool-Calling and Answer Generation
2 benchmarks · 2 papers
Multi-tool calling
2 benchmarks · 1 papers
Tool Sequence Recommendation
2 benchmarks · 1 papers
Tool Use & Coding
Page 13 of 49
Previous
Next