Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 26 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
1 benchmarks · 1 papers
Natural-language music software control
1 benchmarks · 1 papers
Database task execution
1 benchmarks · 1 papers
Multi-modal Agent Reasoning
1 benchmarks · 1 papers
Autonomous Exploration and Discovery
1 benchmarks · 1 papers
Multi-task Agent Execution
1 benchmarks · 1 papers
Adversarial Attack against SeeAct agent
1 benchmarks · 1 papers
Adversarial Attack against WebExperT agent
1 benchmarks · 1 papers
Action Routing
1 benchmarks · 1 papers
Agent Distillation
1 benchmarks · 1 papers
Proactive Agent Task Execution
1 benchmarks · 1 papers
Reward Hacking
1 benchmarks · 1 papers
Trajectory search
1 benchmarks · 1 papers
Long-Horizon World Modelling
1 benchmarks · 1 papers
Agentic Memory Management
1 benchmarks · 1 papers
SSP Planning
1 benchmarks · 1 papers
API Execution Simulation
1 benchmarks · 1 papers
Mobile-use agent task completion and intent alignment
1 benchmarks · 1 papers
OS agent GUI interaction
1 benchmarks · 1 papers
Desktop UI Navigation
1 benchmarks · 2 papers
Competitive Programming Agent Evaluation
1 benchmarks · 1 papers
Compositional planning
1 benchmarks · 1 papers
Behavior Tree Synthesis
1 benchmarks · 2 papers
Windows UI Navigation
1 benchmarks · 1 papers
Agent Toolchain Scheduling
1 benchmarks · 1 papers
Hijacking detection
1 benchmarks · 3 papers
AI Agent Reasoning and Tool-use
1 benchmarks · 1 papers
2-player Diplomacy
1 benchmarks · 1 papers
Web UI Navigation
1 benchmarks · 1 papers
Agent-based Data Analysis
1 benchmarks · 1 papers
Agent Safety Reasoning
1 benchmarks · 1 papers
SQL code generation agent
1 benchmarks · 1 papers
Travel planning agent
Page 26 of 49
Previous
Next