Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 17 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
5 benchmarks · 1 papers
Hybrid Reasoning
5 benchmarks · 1 papers
LLM Judgment
5 benchmarks · 8 papers
General Knowledge and Reasoning
5 benchmarks · 1 papers
Prosocial Alignment
5 benchmarks · 1 papers
Reasoning failure prediction and recovery
5 benchmarks · 2 papers
Social Interaction Evaluation
5 benchmarks · 1 papers
Deliberative Reason Index (DRI) calculation
5 benchmarks · 3 papers
Proactive dialogue
5 benchmarks · 1 papers
Human-Metric Correlation
5 benchmarks · 3 papers
Memory Question Answering
5 benchmarks · 7 papers
Language Reasoning
5 benchmarks · 1 papers
Math and Text Reasoning
5 benchmarks · 1 papers
Text-based Task Completion
5 benchmarks · 1 papers
Qualitative Evaluation of Stance Distribution and Argument Organization
5 benchmarks · 3 papers
Reward alignment
5 benchmarks · 4 papers
LLM Generation
5 benchmarks · 6 papers
Hallucination Robustness
5 benchmarks · 17 papers
Over-refusal
5 benchmarks · 4 papers
Pairwise Preference Evaluation
5 benchmarks · 7 papers
Multi-Turn Tool Calling
5 benchmarks · 5 papers
Multi-turn Dialogue Reasoning
5 benchmarks · 1 papers
Tool Invocation Refusal
5 benchmarks · 1 papers
Sequential Composition Generalization
5 benchmarks · 2 papers
Lifelong Model Editing
5 benchmarks · 5 papers
Zero-shot language evaluation
5 benchmarks · 1 papers
Experience Reuse
5 benchmarks · 4 papers
Sycophancy
5 benchmarks · 1 papers
Legal Contract Revision
5 benchmarks · 1 papers
LLM Filtering
5 benchmarks · 1 papers
Multi-turn Inference Latency
5 benchmarks · 1 papers
Long speech understanding and reasoning
5 benchmarks · 1 papers
Task Vector Performance
Page 17 of 154
Previous
Next