Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 21 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
5 benchmarks · 1 papers
Comprehension Questions
5 benchmarks · 1 papers
Over-refusal Compliance
5 benchmarks · 8 papers
General Knowledge and Reasoning
5 benchmarks · 1 papers
Prosocial Alignment
5 benchmarks · 1 papers
Personalized VLM Alignment
5 benchmarks · 1 papers
Personalized response selection
5 benchmarks · 1 papers
Tool Invocation Refusal
5 benchmarks · 2 papers
Social Interaction Evaluation
5 benchmarks · 1 papers
Black-box LLM Jailbreaking
5 benchmarks · 2 papers
Lifelong Model Editing
5 benchmarks · 1 papers
Legal Contract Revision
5 benchmarks · 1 papers
LLM Judgment
5 benchmarks · 1 papers
Sequential Composition Generalization
5 benchmarks · 1 papers
Process-level Evaluation
5 benchmarks · 1 papers
Human-Metric Correlation
5 benchmarks · 1 papers
Free Q&A
4 benchmarks · 1 papers
Disruption
4 benchmarks · 1 papers
Proactive next utterance prediction
4 benchmarks · 1 papers
CoT Naturalness
4 benchmarks · 2 papers
Long-Horizon Stability
4 benchmarks · 5 papers
Accuracy
4 benchmarks · 1 papers
CoT Soundness Evaluation
4 benchmarks · 1 papers
Multi-party Dialogue Generation
4 benchmarks · 2 papers
Multi-turn Dialogue Response Generation
4 benchmarks · 1 papers
Information Analysis and Processing
4 benchmarks · 1 papers
Knowledge Distillation Robustness
4 benchmarks · 1 papers
Reasoning over Large Structured Context
4 benchmarks · 2 papers
LLM Training Performance
4 benchmarks · 2 papers
Category Unlearning
4 benchmarks · 1 papers
Compositional Steering
4 benchmarks · 5 papers
Open-ended writing
4 benchmarks · 5 papers
Memory Extraction
Page 21 of 154
Previous
Next