Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 9 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
4 benchmarks · 2 papers
Conversation Simulation
4 benchmarks · 2 papers
Sequential Planning
4 benchmarks · 1 papers
Mobile Agent Action Execution
4 benchmarks · 4 papers
GUI Agent Task Success
4 benchmarks · 5 papers
General AI Assistant
4 benchmarks · 1 papers
Graph-based Agent Memory Poisoning
4 benchmarks · 1 papers
Interactive Agent
4 benchmarks · 1 papers
Agentic Tool Calling
4 benchmarks · 1 papers
Multi-Agent Game
4 benchmarks · 1 papers
Agentic Skill Execution
4 benchmarks · 14 papers
General AI Assistant Task
4 benchmarks · 1 papers
Itinerary Planning
3 benchmarks · 1 papers
Video Scrubbing
3 benchmarks · 1 papers
Sequential Crafting
3 benchmarks · 1 papers
Cross-factor predictive validity
3 benchmarks · 1 papers
Lifelong Agent Interaction
3 benchmarks · 1 papers
Action Interception
3 benchmarks · 2 papers
Step-level reasoning evaluation
3 benchmarks · 3 papers
Web Task Execution
3 benchmarks · 3 papers
Agent Planning
3 benchmarks · 1 papers
Agent action prediction
3 benchmarks · 1 papers
Multi-agent Cognitive Orchestration
3 benchmarks · 1 papers
Tool-use Inference
3 benchmarks · 1 papers
Long-horizon procedural planning
3 benchmarks · 1 papers
GUI planning semantic consistency
3 benchmarks · 1 papers
Web Navigation Task
3 benchmarks · 1 papers
Agentic Workflow Automation
3 benchmarks · 1 papers
Web GUI Task Automation
3 benchmarks · 2 papers
Social Deduction Game Gameplay
3 benchmarks · 1 papers
Safe Tool Execution
3 benchmarks · 1 papers
LLM Agent Defense
3 benchmarks · 1 papers
Tool-use API Generalization
Page 9 of 49
Previous
Next