Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 96 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
1 benchmarks · 1 papers
AI-Agent Compatibility Evaluation
1 benchmarks · 1 papers
Stability
1 benchmarks · 1 papers
Theory of Mind Knowledge Prediction
1 benchmarks · 1 papers
Zero-shot Commonsense Reasoning and Knowledge
1 benchmarks · 1 papers
Error propagation analysis
1 benchmarks · 1 papers
Cross-domain language model performance evaluation
1 benchmarks · 1 papers
Zero-shot Evaluation Aggregate
1 benchmarks · 3 papers
Online Inference
1 benchmarks · 1 papers
Agent Selection
1 benchmarks · 1 papers
Large Language Model Training Efficiency
1 benchmarks · 1 papers
World Knowledge and Reading Comprehension
1 benchmarks · 3 papers
Abductive Commonsense Reasoning
1 benchmarks · 1 papers
Theory of Mind - Intention
1 benchmarks · 1 papers
Agent Task-solving
1 benchmarks · 1 papers
Influential training example identification
1 benchmarks · 1 papers
Self-awareness
1 benchmarks · 1 papers
End-to-end task completion
1 benchmarks · 1 papers
Simulated Patient Dialogue Generation
1 benchmarks · 1 papers
Graduate-level Factual Reasoning
1 benchmarks · 1 papers
Open-ended Climate Analysis
1 benchmarks · 1 papers
Long-term memory and reasoning
1 benchmarks · 1 papers
Standardized Patient Simulation
1 benchmarks · 1 papers
Strategic Scenario Quality Assessment
1 benchmarks · 1 papers
General Multimodal Intelligence
1 benchmarks · 1 papers
Physical Paradox Detection
1 benchmarks · 1 papers
Physical reasoning and problem-solving
1 benchmarks · 1 papers
Instruction-following and procedural reasoning
1 benchmarks · 2 papers
Sequential task management and state maintenance
1 benchmarks · 1 papers
General Usability Evaluation
1 benchmarks · 1 papers
Analytical Agent Evaluation
1 benchmarks · 1 papers
Format Debiasing
1 benchmarks · 1 papers
Analytical Agent Reasoning
Page 96 of 154
Previous
Next