Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 35 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
1 benchmarks · 1 papers
Desktop automation
1 benchmarks · 1 papers
Scientific Reasoning in Text-based Environments
1 benchmarks · 1 papers
Multi-Agent Semantic Navigation
1 benchmarks · 1 papers
Tool Drawer Org.
1 benchmarks · 1 papers
Autonomous Agent Performance
1 benchmarks · 1 papers
Generic GUI Execution
1 benchmarks · 1 papers
Agent Action
1 benchmarks · 1 papers
Visual Tool Reasoning
1 benchmarks · 1 papers
Agent Defense
1 benchmarks · 1 papers
Math Reasoning (coding tools)
1 benchmarks · 1 papers
Financial Tool Usage
1 benchmarks · 1 papers
Generative Modeling Efficiency
1 benchmarks · 1 papers
Execution Governance Evaluation
1 benchmarks · 1 papers
Agent Return Volatility
1 benchmarks · 2 papers
Malicious behavior measurement
1 benchmarks · 1 papers
Agentic Rollback and Planning
1 benchmarks · 1 papers
Multi-agent execution
1 benchmarks · 1 papers
Multi-step tool-use composition
1 benchmarks · 1 papers
Workflow Generation and Execution
1 benchmarks · 1 papers
Agentic Benchmarks
1 benchmarks · 1 papers
Benign tool-calling reliability
1 benchmarks · 1 papers
Agentic-Level Attack
1 benchmarks · 1 papers
Multi-agent software engineering coordination
1 benchmarks · 2 papers
Tool-based Manipulation
1 benchmarks · 1 papers
Benign completion reliability
1 benchmarks · 1 papers
Large Language Model Routing and Orchestration
1 benchmarks · 1 papers
GUI Agent Planning
1 benchmarks · 1 papers
Functionality Grounding
1 benchmarks · 3 papers
Omni-modal agentic tasks
1 benchmarks · 1 papers
Autonomous GUI Interaction
1 benchmarks · 1 papers
Agent Drift Mitigation
1 benchmarks · 1 papers
Opponent Exploitation
Page 35 of 49
Previous
Next