Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 57 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
2 benchmarks · 1 papers
Long-horizon dialogue
2 benchmarks · 1 papers
Pluralistic Reward Model Learning
2 benchmarks · 2 papers
Sycophancy Reduction
2 benchmarks · 2 papers
Engineering problem-solving
2 benchmarks · 1 papers
Deception Evaluation
2 benchmarks · 2 papers
General Alignment
2 benchmarks · 1 papers
Open-domain task
2 benchmarks · 1 papers
Sequential Multi-Task Learning
2 benchmarks · 2 papers
Strike
2 benchmarks · 1 papers
Dialogue Commonsense Reasoning
2 benchmarks · 1 papers
Long-Horizon Search Intelligence
2 benchmarks · 1 papers
Interactive query answering
2 benchmarks · 1 papers
Article Writing
2 benchmarks · 2 papers
Tutoring Evaluation
2 benchmarks · 1 papers
Language-based Games
2 benchmarks · 4 papers
Long-context reasoning and retrieval
2 benchmarks · 2 papers
Secure LLM Agent Task Completion
2 benchmarks · 2 papers
Common Sense Reasoning and Question Answering
2 benchmarks · 2 papers
Usability Evaluation
2 benchmarks · 2 papers
Belief Prediction
2 benchmarks · 2 papers
LLM Hallucination Detection
2 benchmarks · 1 papers
Human-AI Agreement Assessment
2 benchmarks · 3 papers
grade-school math
2 benchmarks · 1 papers
Reasoning and Synthesis
2 benchmarks · 1 papers
General Knowledge and Instruction Following
2 benchmarks · 2 papers
Islamic inheritance reasoning
2 benchmarks · 1 papers
Evidence-grounded diagnostic reasoning
2 benchmarks · 1 papers
Defective Dialog Detection
2 benchmarks · 1 papers
Instructed Code Generation
2 benchmarks · 1 papers
Uncovering hidden system prompts
2 benchmarks · 1 papers
LLM-rated generation quality
2 benchmarks · 1 papers
Standard Operating Procedure execution
Page 57 of 154
Previous
Next