Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 7 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
4 benchmarks · 1 papers
Graph-based Agent Memory Poisoning
4 benchmarks · 5 papers
Multimodal Tool Use
4 benchmarks · 1 papers
Tool Retrieval and Function Selection
4 benchmarks · 2 papers
Skill Composition
4 benchmarks · 1 papers
Multi-agent system task solving
4 benchmarks · 5 papers
Web automation
4 benchmarks · 4 papers
Multi-task performance evaluation
4 benchmarks · 5 papers
Long-horizon Task Execution
4 benchmarks · 3 papers
Agent Security
4 benchmarks · 1 papers
Multi-Turn Tool-Integrated Reasoning (TIR)
4 benchmarks · 1 papers
Multi-Agent Game
4 benchmarks · 1 papers
Agentic Skill Execution
4 benchmarks · 1 papers
General Deep Research Tool Use
4 benchmarks · 1 papers
Tool Attribution Correctness
4 benchmarks · 1 papers
Itinerary Planning
4 benchmarks · 1 papers
Tool-use Planning
4 benchmarks · 1 papers
Multi-agent task routing
4 benchmarks · 1 papers
Function Invocation
4 benchmarks · 4 papers
GUI Agent Task Success
4 benchmarks · 1 papers
Multi-agent planning and execution
4 benchmarks · 16 papers
Online Shopping
4 benchmarks · 1 papers
GUI
4 benchmarks · 1 papers
Embodied AI Planning
4 benchmarks · 2 papers
User simulator goal alignment
4 benchmarks · 1 papers
First-try Success
4 benchmarks · 1 papers
Web Action Generation Efficiency
4 benchmarks · 1 papers
Failure Flipping
4 benchmarks · 1 papers
Function Selection
4 benchmarks · 2 papers
Sequential Planning
4 benchmarks · 2 papers
Dialogue Agent Interaction
4 benchmarks · 1 papers
Multi-Turn Constraint Adaptation
4 benchmarks · 1 papers
Tool-use behavioral profiling
Page 7 of 49
Previous
Next