Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 34 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
1 benchmarks · 1 papers
Long-Horizon Tool Execution
1 benchmarks · 1 papers
Multi-agent interaction and social reasoning
1 benchmarks · 1 papers
Distractor Selection
1 benchmarks · 4 papers
GUI Navigation and Action
1 benchmarks · 1 papers
Multi-agent research collaboration
1 benchmarks · 1 papers
Freight Negotiation
1 benchmarks · 1 papers
Guide target element selection
1 benchmarks · 1 papers
Cross-Lingual Planning
1 benchmarks · 1 papers
Multi-Player Social Reasoning and Strategy
1 benchmarks · 1 papers
Active Tool Exploration
1 benchmarks · 1 papers
Indirect Prompt Injection Defense Evaluation
1 benchmarks · 1 papers
Comparative Analysis of System Dimensions
1 benchmarks · 1 papers
Prepare lunch
1 benchmarks · 1 papers
Embodied Visual Agent Task
1 benchmarks · 1 papers
Desktop automation
1 benchmarks · 1 papers
Scientific Reasoning in Text-based Environments
1 benchmarks · 1 papers
Multi-Agent Semantic Navigation
1 benchmarks · 1 papers
Tool Drawer Org.
1 benchmarks · 1 papers
Autonomous Agent Performance
1 benchmarks · 1 papers
In-distribution Tool Use
1 benchmarks · 1 papers
Generic GUI Execution
1 benchmarks · 1 papers
Math Reasoning (coding tools)
1 benchmarks · 1 papers
Financial Tool Usage
1 benchmarks · 1 papers
Execution Governance Evaluation
1 benchmarks · 2 papers
Malicious behavior measurement
1 benchmarks · 1 papers
Agentic Rollback and Planning
1 benchmarks · 1 papers
Multi-agent execution
1 benchmarks · 1 papers
Workflow Generation and Execution
1 benchmarks · 1 papers
Agentic Benchmarks
1 benchmarks · 1 papers
Benign tool-calling reliability
1 benchmarks · 1 papers
Agentic-Level Attack
1 benchmarks · 2 papers
Tool-based Manipulation
Page 34 of 49
Previous
Next