Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 33 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
1 benchmarks · 1 papers
Tool Selection Quality
1 benchmarks · 1 papers
Multi-Agent Edge Computing Orchestration
1 benchmarks · 1 papers
Agent Memory Question Answering
1 benchmarks · 1 papers
Tool orchestration
1 benchmarks · 1 papers
API Calling
1 benchmarks · 1 papers
End-to-end task execution
1 benchmarks · 1 papers
Workflow Planning
1 benchmarks · 1 papers
Atomic Task Execution
1 benchmarks · 1 papers
Multi-step web search
1 benchmarks · 1 papers
Medical Visual Question Answering / Tool-use
1 benchmarks · 1 papers
Atomic Task Success
1 benchmarks · 1 papers
Long-Horizon Tool Execution
1 benchmarks · 1 papers
Multi-agent interaction and social reasoning
1 benchmarks · 1 papers
General AI Assistant Task Completion
1 benchmarks · 1 papers
Multi-agent research collaboration
1 benchmarks · 1 papers
Freight Negotiation
1 benchmarks · 1 papers
Guide target element selection
1 benchmarks · 1 papers
Cross-Lingual Planning
1 benchmarks · 1 papers
Multi-Player Social Reasoning and Strategy
1 benchmarks · 1 papers
Comparative Analysis of System Dimensions
1 benchmarks · 1 papers
Prepare lunch
1 benchmarks · 1 papers
Embodied Visual Agent Task
1 benchmarks · 1 papers
Desktop automation
1 benchmarks · 1 papers
Scientific Reasoning in Text-based Environments
1 benchmarks · 1 papers
Multi-Agent Semantic Navigation
1 benchmarks · 1 papers
Tool Drawer Org.
1 benchmarks · 1 papers
Autonomous Agent Performance
1 benchmarks · 1 papers
Generic GUI Execution
1 benchmarks · 1 papers
Math Reasoning (coding tools)
1 benchmarks · 1 papers
Financial Tool Usage
1 benchmarks · 1 papers
Execution Governance Evaluation
1 benchmarks · 2 papers
Malicious behavior measurement
Page 33 of 49
Previous
Next