Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 46 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
1 benchmarks · 1 papers
Financial Agent Task
1 benchmarks · 1 papers
Multi-Task Agent Generalization
1 benchmarks · 1 papers
Tutor Agent Evaluation
1 benchmarks · 1 papers
Agentic Workflow Injection Detection
1 benchmarks · 1 papers
Multi-agent game strategic reasoning
1 benchmarks · 1 papers
Strategic reasoning in auction games
1 benchmarks · 1 papers
Plan-level Tool Use
1 benchmarks · 1 papers
Multi-constraint search problem solving
1 benchmarks · 1 papers
Household Task Execution
1 benchmarks · 1 papers
Science Experiment Execution
1 benchmarks · 1 papers
Agentic Automation
1 benchmarks · 1 papers
Multi-Agent Pickup and Delivery Planning
1 benchmarks · 1 papers
puzzle-4x4-task4
1 benchmarks · 1 papers
Meeting Scheduling
1 benchmarks · 1 papers
Agentic Compositional Generalization
1 benchmarks · 1 papers
Sequential Tool Use
1 benchmarks · 1 papers
Software Engineering Agent Task
1 benchmarks · 1 papers
Online Shopping and Web Navigation
1 benchmarks · 1 papers
Tool Activation Probing
1 benchmarks · 1 papers
UI Task Completion
1 benchmarks · 1 papers
Stateful Agent-User Interaction
1 benchmarks · 1 papers
HYBRID
1 benchmarks · 1 papers
Enterprise Task Execution
1 benchmarks · 1 papers
Multi-agent consensus coordination
1 benchmarks · 1 papers
Knowledge-based Agent Reasoning
1 benchmarks · 1 papers
End-to-end answer accuracy
1 benchmarks · 1 papers
Web Browsing Task
1 benchmarks · 1 papers
Agent Perturbation Reliability Testing
1 benchmarks · 1 papers
Multi-agent task coordination
1 benchmarks · 1 papers
Compound LLM Collaboration
1 benchmarks · 1 papers
Reasoning-Level Denial-of-Service Attack
1 benchmarks · 1 papers
Operating System GUI Agentic Reasoning
Page 46 of 49
Previous
Next