Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 42 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
1 benchmarks · 1 papers
Agent Retrieval
1 benchmarks · 1 papers
App-based Agentic Task
1 benchmarks · 1 papers
Opponent modeling
1 benchmarks · 1 papers
Skill invocation in LLM agents
1 benchmarks · 1 papers
Negotiation performance and belief calibration
1 benchmarks · 1 papers
Safe Agent Evaluation
1 benchmarks · 1 papers
Negotiation task (Sell&Buy)
1 benchmarks · 1 papers
Agent Action Safety Verification
1 benchmarks · 1 papers
Tool Retrieval and Invocation
1 benchmarks · 1 papers
Language Model Routing and Orchestration
1 benchmarks · 1 papers
Windows Agent Task Completion
1 benchmarks · 1 papers
Terminal-based Tool Use
1 benchmarks · 1 papers
Stateful Agent Backdoor Attack Evaluation
1 benchmarks · 1 papers
Operating System Agent Task Completion
1 benchmarks · 1 papers
Stateful Agent Backdoor Attack Performance
1 benchmarks · 1 papers
Benign Task Completion
1 benchmarks · 1 papers
Language-based Planning
1 benchmarks · 1 papers
Stateful Agent Backdoor Attack
1 benchmarks · 1 papers
Windows Computer Use Automation
1 benchmarks · 1 papers
Web-search agent tasks
1 benchmarks · 1 papers
Safe Agent Planning
1 benchmarks · 1 papers
Agent Tool-use and Reasoning
1 benchmarks · 1 papers
GUI State Control
1 benchmarks · 1 papers
Code-interpreter tasks
1 benchmarks · 1 papers
Earth Observation Agent Task Completion
1 benchmarks · 1 papers
LLM Agent Safety
1 benchmarks · 1 papers
Tool-use performance
1 benchmarks · 1 papers
Text-based embodied task completion
1 benchmarks · 3 papers
Tool Use and Reasoning
1 benchmarks · 1 papers
Task performance (Cart Score)
1 benchmarks · 1 papers
Workflow Execution (Average)
1 benchmarks · 1 papers
Agentic Workflow Performance (Iterative Refinement Loops)
Page 42 of 49
Previous
Next