Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 17 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
2 benchmarks · 1 papers
Multi-step reasoning and knowledge retrieval
2 benchmarks · 1 papers
Inquire-driven GUI Navigation
2 benchmarks · 1 papers
Multi-agent Semantic Object Navigation
2 benchmarks · 1 papers
Safety and Action Blocking
2 benchmarks · 3 papers
GUI Task Completion
2 benchmarks · 2 papers
Interactive tool-use
2 benchmarks · 1 papers
Single-agent contract design
2 benchmarks · 1 papers
Multi-agent contract design
2 benchmarks · 1 papers
Steering Validation
2 benchmarks · 1 papers
LLM Agent Security and Utility Evaluation
2 benchmarks · 1 papers
Enterprise interface interaction
2 benchmarks · 3 papers
Terminal-based task execution
2 benchmarks · 1 papers
Enterprise interface task completion
2 benchmarks · 1 papers
Knowledge-to-action question answering
2 benchmarks · 2 papers
Computer task execution
2 benchmarks · 2 papers
General Agent Capability
2 benchmarks · 2 papers
Agent Routing
2 benchmarks · 1 papers
Driving Hazard Explanation and Action Generation
2 benchmarks · 18 papers
Web Navigation and Shopping
2 benchmarks · 1 papers
Human-agent alignment
2 benchmarks · 5 papers
Mobile UI Control
2 benchmarks · 1 papers
Form-filling
2 benchmarks · 1 papers
device control
2 benchmarks · 1 papers
Agentic medical interaction
2 benchmarks · 1 papers
Tool-use agentic performance
2 benchmarks · 1 papers
GUI Agent Automation
2 benchmarks · 1 papers
Regions of Interest Discovery
2 benchmarks · 1 papers
Agentic Affordance Reasoning
2 benchmarks · 1 papers
Mobile Agent Interaction
2 benchmarks · 4 papers
Agentic Web Browsing
2 benchmarks · 1 papers
Offline Constrained RLHF
2 benchmarks · 1 papers
Agent Planning and API Calling
Page 17 of 49
Previous
Next