Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
Agents & Tool Use AI research – Page 21 · SOTA2 Research
Back to Research
Research domain
Agents & Tool Use
Search tasks
Search
1 benchmarks · 1 papers
Multi-turn Strategic Gameplay
1 benchmarks · 1 papers
Diplomacy
1 benchmarks · 1 papers
Web Browsing Competition
1 benchmarks · 1 papers
Interactive agent tasks
1 benchmarks · 1 papers
Tool-use AI Assistant
1 benchmarks · 1 papers
Tool-selection bias evaluation
1 benchmarks · 2 papers
Web Browsing Competition (Chinese)
1 benchmarks · 1 papers
Agent Generalization
1 benchmarks · 1 papers
Multi-agent policy generation
1 benchmarks · 4 papers
GUI Navigation and Action
1 benchmarks · 1 papers
End-to-end terminal tasks
1 benchmarks · 1 papers
Agentic Multimodal Tool-use
1 benchmarks · 4 papers
Task and Scenario Goal Completion
1 benchmarks · 1 papers
Code Agent Simulation
1 benchmarks · 1 papers
Indirect Prompt Injection Defense Evaluation
1 benchmarks · 1 papers
Search Agent Reasoning
1 benchmarks · 1 papers
API Invocation Completion
1 benchmarks · 2 papers
Multimodal Reasoning and Tool-use
1 benchmarks · 1 papers
Mandate Tool Invocation
1 benchmarks · 1 papers
Sole Planning
1 benchmarks · 1 papers
In-distribution Tool Use
1 benchmarks · 1 papers
Planning-style reasoning
1 benchmarks · 1 papers
Financial Specialist Tool Use
1 benchmarks · 1 papers
Software Engineering Tool Use
1 benchmarks · 1 papers
Zero-Shot Generalist Tool Use
1 benchmarks · 1 papers
Agent objective detection
1 benchmarks · 1 papers
Routine Task Management
1 benchmarks · 1 papers
Agent Defense
1 benchmarks · 1 papers
Visual Tool Reasoning
1 benchmarks · 1 papers
Personalized Assistant Interaction
1 benchmarks · 1 papers
Mobile Application Operation
1 benchmarks · 1 papers
Language-based decision-making
Page 21 of 49
Previous
Next