Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 41 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
1 benchmarks · 1 papers
Tool-use Agent Task
1 benchmarks · 1 papers
Multimodal Healthcare Agent Performance
1 benchmarks · 1 papers
Multimedia Plan Execution (MI-T)
1 benchmarks · 1 papers
Language Agent Memory Management
1 benchmarks · 1 papers
Multi-agent pursuit-evasion
1 benchmarks · 1 papers
Automated Research
1 benchmarks · 1 papers
Agentic Mobile Interaction
1 benchmarks · 1 papers
Agent Retrieval
1 benchmarks · 1 papers
App-based Agentic Task
1 benchmarks · 1 papers
Agentic Memory Retrieval
1 benchmarks · 1 papers
Multimedia Plan Execution (MV-A)
1 benchmarks · 1 papers
Process reward modeling evaluation
1 benchmarks · 1 papers
Multi-Step Tool Orchestration
1 benchmarks · 1 papers
GUI Automation Replay
1 benchmarks · 1 papers
Earth Observation agent reasoning
1 benchmarks · 1 papers
Multimedia Plan Execution (MV-T)
1 benchmarks · 1 papers
Opponent modeling
1 benchmarks · 1 papers
Skill invocation in LLM agents
1 benchmarks · 1 papers
Negotiation performance and belief calibration
1 benchmarks · 1 papers
Safe Agent Evaluation
1 benchmarks · 1 papers
Negotiation task (Sell&Buy)
1 benchmarks · 1 papers
Agent Action Safety Verification
1 benchmarks · 1 papers
Language Model Routing and Orchestration
1 benchmarks · 1 papers
Windows Agent Task Completion
1 benchmarks · 1 papers
Terminal-based Tool Use
1 benchmarks · 1 papers
Stateful Agent Backdoor Attack Evaluation
1 benchmarks · 1 papers
Operating System Agent Task Completion
1 benchmarks · 1 papers
Stateful Agent Backdoor Attack Performance
1 benchmarks · 1 papers
Stateful Agent Backdoor Attack
1 benchmarks · 1 papers
Windows Computer Use Automation
1 benchmarks · 1 papers
Web-search agent tasks
1 benchmarks · 1 papers
Safe Agent Planning
Page 41 of 49
Previous
Next