Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 8 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
4 benchmarks · 1 papers
Failure Flipping
4 benchmarks · 2 papers
Dialogue Agent Interaction
4 benchmarks · 1 papers
Multi-agent system task solving
4 benchmarks · 1 papers
Function Selection
4 benchmarks · 1 papers
Graph-based Agent Memory Poisoning
4 benchmarks · 1 papers
Tool Retrieval and Function Selection
4 benchmarks · 1 papers
Multi-Turn Tool-Integrated Reasoning (TIR)
4 benchmarks · 3 papers
GUI task automation
4 benchmarks · 1 papers
Tool Attribution Correctness
4 benchmarks · 4 papers
GUI Interaction
4 benchmarks · 2 papers
Sequential Planning
4 benchmarks · 1 papers
Function Invocation
4 benchmarks · 1 papers
Interactive Agent
4 benchmarks · 1 papers
Multi-Agent Game
4 benchmarks · 4 papers
GUI Agent Task Success
4 benchmarks · 2 papers
Computer-Using Agent Task
4 benchmarks · 1 papers
Prompting algorithm implementation
4 benchmarks · 1 papers
Agentic Skill Execution
4 benchmarks · 2 papers
Agentic Evaluation
4 benchmarks · 1 papers
Temporal Numeric Planning
4 benchmarks · 1 papers
GUI
4 benchmarks · 1 papers
Embodied AI Planning
4 benchmarks · 1 papers
Preference-driven Tool Calling
4 benchmarks · 1 papers
Guarded Agent Evaluation
4 benchmarks · 3 papers
Agent Security
4 benchmarks · 1 papers
Multi-agent edge orchestration
4 benchmarks · 1 papers
General Deep Research Tool Use
4 benchmarks · 1 papers
Hybrid planning refinement
4 benchmarks · 1 papers
Tool-use behavioral profiling
4 benchmarks · 1 papers
Multi-agent planning and execution
4 benchmarks · 5 papers
General AI Assistant
4 benchmarks · 1 papers
LLM-Agent Oversight
Page 8 of 49
Previous
Next