Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 24 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
1 benchmarks · 1 papers
Mobile operating system task execution
1 benchmarks · 1 papers
Grounded Chess Reasoning
1 benchmarks · 1 papers
Retrieve Item from Safe
1 benchmarks · 1 papers
Desktop operating system task execution
1 benchmarks · 1 papers
Multi-stage agent workflow execution
1 benchmarks · 1 papers
Scientific Agent Task
1 benchmarks · 1 papers
Agentic cryptocurrency trading
1 benchmarks · 1 papers
Autonomous Defense
1 benchmarks · 1 papers
Native Windows operating system task execution
1 benchmarks · 1 papers
Multimedia Plan Execution (AV-A)
1 benchmarks · 1 papers
Task Completion Rate
1 benchmarks · 1 papers
Multimedia Plan Execution (AV-T)
1 benchmarks · 1 papers
Enterprise task completion
1 benchmarks · 1 papers
Tool-using Reasoning
1 benchmarks · 1 papers
Multimedia Plan Execution (IA-T)
1 benchmarks · 1 papers
Tool-use Agent Robustness
1 benchmarks · 1 papers
Multimedia Plan Execution (IA-V)
1 benchmarks · 1 papers
Agent defense evaluation
1 benchmarks · 1 papers
Multimedia Plan Execution (IV-A)
1 benchmarks · 1 papers
Earth Observation Agent Performance
1 benchmarks · 1 papers
Multimedia Plan Execution (IV-T)
1 benchmarks · 1 papers
Cross-server tool sequence generation
1 benchmarks · 1 papers
Multimedia Plan Execution (IV-V)
1 benchmarks · 1 papers
Shared Autonomy
1 benchmarks · 1 papers
Multimedia Plan Execution (MA-I)
1 benchmarks · 1 papers
LLM Agent Security Defense
1 benchmarks · 1 papers
Multimedia Plan Execution (MA-T)
1 benchmarks · 1 papers
Agentic UI Interaction
1 benchmarks · 1 papers
Multimedia Plan Execution (MA-V)
1 benchmarks · 1 papers
Mobile GUI Agent Backdoor Defense
1 benchmarks · 1 papers
Multimedia Plan Execution (MI-A)
1 benchmarks · 1 papers
One-step plan validity
Page 24 of 49
Previous
Next