Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 40 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
1 benchmarks · 1 papers
Agent Adaptation
1 benchmarks · 1 papers
Multimodal Agentic Tool Use
1 benchmarks · 1 papers
Multimedia Plan Execution (IA-V)
1 benchmarks · 1 papers
clean_up_your_desk
1 benchmarks · 1 papers
Due Diligence
1 benchmarks · 1 papers
Workflow Discovery
1 benchmarks · 1 papers
Autonomous Cybersecurity Defense
1 benchmarks · 1 papers
Tool-use interaction evaluation
1 benchmarks · 1 papers
Financial Advisory Copilot
1 benchmarks · 1 papers
Sequential Execution Validation
1 benchmarks · 1 papers
Smart home command understanding
1 benchmarks · 1 papers
Multi-Agent Orchestration
1 benchmarks · 1 papers
Multimedia Plan Execution (IV-T)
1 benchmarks · 1 papers
Agentic Time Series Analysis
1 benchmarks · 1 papers
Shared Autonomy
1 benchmarks · 1 papers
Long-horizon LLM Agent performance
1 benchmarks · 1 papers
Tool-use alignment
1 benchmarks · 1 papers
Safeguarding LLM Agents against prompt injection
1 benchmarks · 1 papers
Coding/Agentic Execution
1 benchmarks · 1 papers
General Artificial Intelligence Capabilities
1 benchmarks · 1 papers
Agent Skill Execution
1 benchmarks · 1 papers
Agentic Reasoning and Interaction
1 benchmarks · 1 papers
LLM Agent Defense Evaluation
1 benchmarks · 1 papers
Group Booking with failures (Grand Rollback)
1 benchmarks · 1 papers
Agentic Mobile Interaction
1 benchmarks · 1 papers
Agent Retrieval
1 benchmarks · 1 papers
App-based Agentic Task
1 benchmarks · 1 papers
Opponent modeling
1 benchmarks · 1 papers
Skill invocation in LLM agents
1 benchmarks · 1 papers
Negotiation performance and belief calibration
1 benchmarks · 1 papers
Safe Agent Evaluation
1 benchmarks · 1 papers
Negotiation task (Sell&Buy)
Page 40 of 49
Previous
Next