Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 39 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
1 benchmarks · 1 papers
UI Navigation / Task Completion
1 benchmarks · 1 papers
Environment-Intensive Task Generation
1 benchmarks · 1 papers
Agentic (multi-turn) evaluation
1 benchmarks · 1 papers
Tool Learning under Instructions Beyond Tool Capabilities
1 benchmarks · 1 papers
Biological Tool Use
1 benchmarks · 1 papers
Desktop task completion
1 benchmarks · 1 papers
Web Browsing and Comparison
1 benchmarks · 1 papers
Web Browsing Automation
1 benchmarks · 1 papers
Proactive Autonomy
1 benchmarks · 1 papers
Scientific Agent Task
1 benchmarks · 3 papers
Agentic Workflow Success
1 benchmarks · 1 papers
Multitask Thai Language Evaluation
1 benchmarks · 1 papers
Meteorological diagnosis agentic workflow
1 benchmarks · 1 papers
Long-horizon item crafting
1 benchmarks · 1 papers
Bimanual tool-use (Scoop into bowl)
1 benchmarks · 1 papers
Single-turn Tool Use
1 benchmarks · 1 papers
Agent Adaptation
1 benchmarks · 1 papers
Due Diligence
1 benchmarks · 1 papers
Workflow Discovery
1 benchmarks · 1 papers
Autonomous Cybersecurity Defense
1 benchmarks · 1 papers
Tool-use interaction evaluation
1 benchmarks · 1 papers
Financial Advisory Copilot
1 benchmarks · 1 papers
Sequential Execution Validation
1 benchmarks · 1 papers
Tool-use alignment
1 benchmarks · 1 papers
Safeguarding LLM Agents against prompt injection
1 benchmarks · 1 papers
Coding/Agentic Execution
1 benchmarks · 1 papers
General Artificial Intelligence Capabilities
1 benchmarks · 1 papers
Agent Skill Execution
1 benchmarks · 1 papers
Agentic Reasoning and Interaction
1 benchmarks · 1 papers
LLM Agent Defense Evaluation
1 benchmarks · 1 papers
Group Booking with failures (Grand Rollback)
1 benchmarks · 1 papers
Agentic Mobile Interaction
Page 39 of 49
Previous
Next