Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 11 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
3 benchmarks · 1 papers
Agent action prediction
3 benchmarks · 1 papers
Skill Learning
3 benchmarks · 1 papers
Next-Action
3 benchmarks · 1 papers
Head-to-Head Evaluation
3 benchmarks · 1 papers
Video Scrubbing
3 benchmarks · 1 papers
Web GUI Task Automation
3 benchmarks · 1 papers
Tool-use API Generalization
3 benchmarks · 1 papers
Program-guided task execution
3 benchmarks · 1 papers
Sequential Crafting
3 benchmarks · 1 papers
Cross-factor predictive validity
3 benchmarks · 10 papers
Long-horizon task completion
3 benchmarks · 1 papers
Lifelong Agent Interaction
3 benchmarks · 2 papers
Failure Recovery
3 benchmarks · 1 papers
Agent Verification
3 benchmarks · 2 papers
Step-level reasoning evaluation
3 benchmarks · 1 papers
Task Planning and Program Generation
3 benchmarks · 4 papers
Multi-Turn Function Calling
3 benchmarks · 1 papers
Tool-use Inference
3 benchmarks · 1 papers
Task-free exploration
3 benchmarks · 1 papers
Action Interception
3 benchmarks · 1 papers
Multi-agent Cognitive Orchestration
3 benchmarks · 1 papers
Web Navigation Task Completion
3 benchmarks · 1 papers
Agentic execution trace analysis
3 benchmarks · 1 papers
Runtime Agent Memory
3 benchmarks · 1 papers
Repeated Negotiation
3 benchmarks · 1 papers
Risk Scenario Generation
3 benchmarks · 1 papers
Agent Trajectory Performance
3 benchmarks · 2 papers
Agent Memory Retrieval
3 benchmarks · 3 papers
Downstream evaluation
3 benchmarks · 2 papers
Terminal Agentic Trajectory Generation
3 benchmarks · 3 papers
Mobile GUI Task Execution
3 benchmarks · 1 papers
Web Navigation Task
Page 11 of 49
Previous
Next