Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 28 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
1 benchmarks · 1 papers
Malicious AI Agent Skill Detection
1 benchmarks · 1 papers
multi-round investment game
1 benchmarks · 1 papers
Agent Success Rate
1 benchmarks · 1 papers
Web Browsing Reasoning
1 benchmarks · 1 papers
Subgoal Planning
1 benchmarks · 2 papers
Web-based interaction
1 benchmarks · 1 papers
Crafting ace items
1 benchmarks · 1 papers
Web navigation / Agent interaction
1 benchmarks · 1 papers
Visual web navigation / Agent interaction
1 benchmarks · 1 papers
Multi-agent steering
1 benchmarks · 1 papers
Semantic Scheduling
1 benchmarks · 1 papers
Tool-augmented Graph Reasoning
1 benchmarks · 1 papers
Injection Attack Detection
1 benchmarks · 1 papers
General Assistant Task Solving
1 benchmarks · 1 papers
Mobile Agent Safety and Capability Evaluation
1 benchmarks · 1 papers
Social Deduction Game Agent Evaluation
1 benchmarks · 1 papers
Operating System Task Automation
1 benchmarks · 1 papers
App management and installation
1 benchmarks · 1 papers
Media playback and ad control
1 benchmarks · 2 papers
Operating System Agent Control
1 benchmarks · 1 papers
Social media interaction
1 benchmarks · 1 papers
Tool usage in multi-turn dialogue
1 benchmarks · 1 papers
Agentic Prompt Injection Defense
1 benchmarks · 2 papers
OS GUI Agentic Task Execution
1 benchmarks · 1 papers
Cross-app workflow
1 benchmarks · 1 papers
Agent Communication Language Evaluation
1 benchmarks · 1 papers
Self-evolution
1 benchmarks · 1 papers
Desktop Application GUI Agent
1 benchmarks · 1 papers
Multi-agent active information gathering
1 benchmarks · 1 papers
Negotiation against a gpt-5.4-high-reasoning seller
1 benchmarks · 1 papers
Autonomous LLM Agent Verification
1 benchmarks · 1 papers
Agent-based interactive task execution
Page 28 of 49
Previous
Next