Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 48 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
1 benchmarks · 1 papers
GUI Agent Execution
1 benchmarks · 1 papers
General Agent
1 benchmarks · 1 papers
Agent Harm Evaluation
1 benchmarks · 1 papers
Cost-aware LLM Routing
1 benchmarks · 1 papers
Cross-app workflow
1 benchmarks · 3 papers
Tool-augmented agent execution
1 benchmarks · 1 papers
Shell-based agent task completion
1 benchmarks · 1 papers
Agent Communication Language Evaluation
1 benchmarks · 1 papers
LLM Agent Memory Retrieval
1 benchmarks · 2 papers
Humanoid Object Search
1 benchmarks · 1 papers
Causal Tool Use
1 benchmarks · 1 papers
Admission Integrity
1 benchmarks · 1 papers
Desktop Application GUI Agent
1 benchmarks · 1 papers
Causal Tool Verification
1 benchmarks · 2 papers
Medical Agent Task Execution
1 benchmarks · 1 papers
Skill Generation
1 benchmarks · 1 papers
Deliberative-reasoning verification
1 benchmarks · 1 papers
Single-hop Tool Calling
1 benchmarks · 1 papers
Agentic Uncertainty Elicitation
1 benchmarks · 1 papers
Web Browsing and Tool Use
1 benchmarks · 1 papers
Agentic Planning for Image Styling
1 benchmarks · 1 papers
Multi-hop tool-use evaluation
1 benchmarks · 1 papers
Embodied Agent Planning
1 benchmarks · 1 papers
Embodied Agent Planning (Safety Evaluation)
1 benchmarks · 1 papers
E-commerce risk management multi-step trajectory evaluation
1 benchmarks · 1 papers
Tool Generation
1 benchmarks · 1 papers
GUI Action Step Prediction
1 benchmarks · 1 papers
Reward Verification
1 benchmarks · 1 papers
Embodied Agent Planning (Adversarial Safety Evaluation)
1 benchmarks · 1 papers
Tool-based multi-turn dialogue
1 benchmarks · 1 papers
UI Operations
1 benchmarks · 1 papers
Verification of Digital Agent Trajectories
Page 48 of 49
Previous
Next