Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 4 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
7 benchmarks · 1 papers
Agent Safety Judgment
7 benchmarks · 6 papers
Web task automation
6 benchmarks · 1 papers
Counterfactual Regret Minimization
6 benchmarks · 1 papers
Autonomous Negotiation
6 benchmarks · 1 papers
Text-based agent interaction
6 benchmarks · 1 papers
Long-horizon agentic tasks
6 benchmarks · 6 papers
GUI Agent Task Completion
6 benchmarks · 4 papers
Agent Tool Use
6 benchmarks · 2 papers
Agentic Capability
6 benchmarks · 1 papers
Agent Authorization Enforcement
6 benchmarks · 1 papers
Grounded Task Planning
6 benchmarks · 1 papers
Instruction Injection Attack on Web Browser Agent
6 benchmarks · 1 papers
Differentiable Planning
6 benchmarks · 1 papers
Preference-conditioned planning
6 benchmarks · 1 papers
Multiagent offline planning
6 benchmarks · 1 papers
Agent Classification
6 benchmarks · 1 papers
Reasoning-Level Denial-of-Service
6 benchmarks · 1 papers
Single-agent tool selection
6 benchmarks · 3 papers
GUI Interaction Control
6 benchmarks · 4 papers
GUI Agent Navigation
6 benchmarks · 1 papers
Agentic Routing
6 benchmarks · 1 papers
Cooperative Multi-agent Problem Solving
6 benchmarks · 1 papers
Long-horizon tasks
6 benchmarks · 2 papers
Terminal-related CLI agent task
6 benchmarks · 4 papers
Tool-use task completion
6 benchmarks · 2 papers
Cybersecurity Challenge Solving
6 benchmarks · 2 papers
Long-horizon agentic task
6 benchmarks · 2 papers
Agent Interaction
6 benchmarks · 1 papers
Agentic Workflow Performance Prediction
6 benchmarks · 1 papers
Multi-Agent Inventory Management
6 benchmarks · 1 papers
Web Agent Automation
6 benchmarks · 4 papers
Multi-turn Agent Interaction
Page 4 of 49
Previous
Next