Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 22 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
1 benchmarks · 1 papers
Agentic Workflow Reconstruction
1 benchmarks · 1 papers
Multi-turn protocol behavior evaluation
1 benchmarks · 2 papers
Multi-agent item delivery
1 benchmarks · 1 papers
Long-Horizon Evaluation with Simulator Feedback
1 benchmarks · 1 papers
Multi-Agent Governance
1 benchmarks · 1 papers
VLM High-level Planning
1 benchmarks · 1 papers
Agentic Orchestration
1 benchmarks · 1 papers
Single-agent tool use
1 benchmarks · 1 papers
Multi-step reasoning and tool use
1 benchmarks · 1 papers
User trait simulation in agentic models
1 benchmarks · 1 papers
Multi-agent code generation defense against sleeper agents
1 benchmarks · 1 papers
Web Navigation Agentic Reasoning
1 benchmarks · 3 papers
Online web agent task success rate
1 benchmarks · 1 papers
Adaptive Tool Use
1 benchmarks · 1 papers
Agent Access Control Evaluation
1 benchmarks · 1 papers
Chinese Web Navigation Agentic Reasoning
1 benchmarks · 1 papers
UI Control
1 benchmarks · 1 papers
Task-Oriented Multi-Agent Communication
1 benchmarks · 1 papers
Tool-agent system evaluation
1 benchmarks · 1 papers
Tool-agent security evaluation
1 benchmarks · 1 papers
LLM Workflow Orchestration
1 benchmarks · 1 papers
Agent action generation
1 benchmarks · 1 papers
Web Agent Attack Success Rate
1 benchmarks · 1 papers
Agentic Presentation Generation
1 benchmarks · 1 papers
Autonomous Cyber Operations
1 benchmarks · 1 papers
Autonomous Purchasing
1 benchmarks · 1 papers
Tool Misuse
1 benchmarks · 1 papers
Autonomous research pipeline execution
1 benchmarks · 1 papers
Tool-use agent security evaluation
1 benchmarks · 1 papers
Injection Task Success
1 benchmarks · 1 papers
Tool Hook Tool Using
1 benchmarks · 1 papers
Tool Pusher Tool Using
Page 22 of 49
Previous
Next