Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 47 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
1 benchmarks · 1 papers
Agentic Compositional Generalization
1 benchmarks · 1 papers
Sequential Tool Use
1 benchmarks · 1 papers
Software Engineering Agent Task
1 benchmarks · 1 papers
Web navigation / Agent interaction
1 benchmarks · 1 papers
Tool Activation Probing
1 benchmarks · 1 papers
UI Task Completion
1 benchmarks · 1 papers
Stateful Agent-User Interaction
1 benchmarks · 1 papers
HYBRID
1 benchmarks · 1 papers
Enterprise Task Execution
1 benchmarks · 1 papers
Multi-agent consensus coordination
1 benchmarks · 1 papers
Knowledge-based Agent Reasoning
1 benchmarks · 1 papers
End-to-end answer accuracy
1 benchmarks · 1 papers
Web Browsing Task
1 benchmarks · 1 papers
Agent Perturbation Reliability Testing
1 benchmarks · 1 papers
Multi-agent task coordination
1 benchmarks · 1 papers
Compound LLM Collaboration
1 benchmarks · 1 papers
Reasoning-Level Denial-of-Service Attack
1 benchmarks · 1 papers
Operating System GUI Agentic Reasoning
1 benchmarks · 1 papers
GUI Agent Execution
1 benchmarks · 1 papers
General Agent
1 benchmarks · 1 papers
Agent Harm Evaluation
1 benchmarks · 1 papers
Cost-aware LLM Routing
1 benchmarks · 3 papers
Tool-augmented agent execution
1 benchmarks · 1 papers
LLM Agent Memory Retrieval
1 benchmarks · 2 papers
Humanoid Object Search
1 benchmarks · 1 papers
Causal Tool Use
1 benchmarks · 1 papers
Causal Tool Verification
1 benchmarks · 2 papers
Medical Agent Task Execution
1 benchmarks · 1 papers
Skill Generation
1 benchmarks · 1 papers
Single-hop Tool Calling
1 benchmarks · 1 papers
Agentic Uncertainty Elicitation
1 benchmarks · 1 papers
Web Browsing and Tool Use
Page 47 of 49
Previous
Next