Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 45 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
1 benchmarks · 1 papers
Tool-call Monitoring
1 benchmarks · 1 papers
Agentic Serving
1 benchmarks · 1 papers
Interactive planning and content generation assistance
1 benchmarks · 1 papers
Web agent meta-evaluation
1 benchmarks · 1 papers
Tool-use runtime evaluation
1 benchmarks · 1 papers
Autonomous Incident Management
1 benchmarks · 1 papers
Terminal task resolution
1 benchmarks · 1 papers
Operating System Operations
1 benchmarks · 1 papers
Web-based shopping
1 benchmarks · 1 papers
Downstream skill execution
1 benchmarks · 1 papers
Web Browsing Action Prediction
1 benchmarks · 1 papers
Single-Agent Household Planning
1 benchmarks · 1 papers
Multi-turn planning
1 benchmarks · 1 papers
Single-Agent Web Navigation
1 benchmarks · 1 papers
Query routing and tool-calling accuracy evaluation
1 benchmarks · 1 papers
Regime Change Detection and Adaptation
1 benchmarks · 1 papers
Agentic Workflow Injection Detection
1 benchmarks · 1 papers
Multi-agent game strategic reasoning
1 benchmarks · 1 papers
E-commerce agentic decision-making
1 benchmarks · 1 papers
End-to-end task success
1 benchmarks · 1 papers
Strategic reasoning in auction games
1 benchmarks · 1 papers
Plan-level Tool Use
1 benchmarks · 1 papers
Multi-constraint search problem solving
1 benchmarks · 1 papers
Household Task Execution
1 benchmarks · 1 papers
Science Experiment Execution
1 benchmarks · 1 papers
puzzle-4x4-task4
1 benchmarks · 1 papers
Agentic Compositional Generalization
1 benchmarks · 1 papers
Sequential Tool Use
1 benchmarks · 1 papers
Software Engineering Agent Task
1 benchmarks · 1 papers
Tool Activation Probing
1 benchmarks · 1 papers
UI Task Completion
1 benchmarks · 1 papers
Stateful Agent-User Interaction
Page 45 of 49
Previous
Next