Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 10 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
3 benchmarks · 1 papers
Next-Action
3 benchmarks · 1 papers
Tool-use Factuality Evaluation
3 benchmarks · 1 papers
LLM Agent Defense
3 benchmarks · 1 papers
Head-to-Head Evaluation
3 benchmarks · 1 papers
Multi-agent system security evaluation
3 benchmarks · 1 papers
Video Scrubbing
3 benchmarks · 1 papers
Embodied Agentic
3 benchmarks · 3 papers
Web Task Execution
3 benchmarks · 1 papers
Agentic Workflow Automation
3 benchmarks · 2 papers
GUI Tasks
3 benchmarks · 1 papers
Tool-augmented agentic reasoning
3 benchmarks · 1 papers
Long-horizon procedural planning
3 benchmarks · 1 papers
Sequential Crafting
3 benchmarks · 1 papers
Agent action prediction
3 benchmarks · 1 papers
Lifelong Agent Interaction
3 benchmarks · 1 papers
Web GUI Task Automation
3 benchmarks · 1 papers
Multi-agent Planning and Composition
3 benchmarks · 1 papers
Cross-factor predictive validity
3 benchmarks · 16 papers
Web navigation and task completion
3 benchmarks · 2 papers
Mobile Agent Evaluation
3 benchmarks · 1 papers
Action Interception
3 benchmarks · 2 papers
Step-level reasoning evaluation
3 benchmarks · 1 papers
Tool-use API Generalization
3 benchmarks · 9 papers
Mobile Task Automation
3 benchmarks · 3 papers
Enterprise workflow automation
3 benchmarks · 3 papers
Web Browsing and Navigation
3 benchmarks · 1 papers
Multi-agent Cognitive Orchestration
3 benchmarks · 1 papers
Tool-use Inference
3 benchmarks · 3 papers
GUI Task Execution
3 benchmarks · 1 papers
Task Planning and Program Generation
3 benchmarks · 1 papers
Safe Tool Execution
3 benchmarks · 2 papers
Social Deduction Game Gameplay
Page 10 of 49
Previous
Next