Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 13 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
7 benchmarks · 11 papers
Long-context language modeling evaluation
7 benchmarks · 11 papers
Multimodal Instruction Following
7 benchmarks · 8 papers
Question Answering and Commonsense Reasoning
7 benchmarks · 3 papers
Multi-subject customization
6 benchmarks · 4 papers
Medical Dialogue Generation
6 benchmarks · 5 papers
Dialog Generation
6 benchmarks · 1 papers
Persona Bias Evaluation
6 benchmarks · 5 papers
Human alignment evaluation
6 benchmarks · 1 papers
Controllability Robustness Prediction
6 benchmarks · 2 papers
LLM Performance Estimation
6 benchmarks · 1 papers
Personalized Long-Form Generation
6 benchmarks · 2 papers
Motivational Interviewing Dialogue Evaluation
6 benchmarks · 1 papers
Long-horizon agentic tasks
6 benchmarks · 4 papers
Long-context Summarization
6 benchmarks · 3 papers
Question Answering with Abstention
6 benchmarks · 7 papers
Multi-Task Reasoning
6 benchmarks · 6 papers
Text Alignment
6 benchmarks · 2 papers
Generative Performance
6 benchmarks · 4 papers
Utility
6 benchmarks · 4 papers
Presentation Generation
6 benchmarks · 3 papers
Speech Instruction-Following
6 benchmarks · 7 papers
Multi-task Language Evaluation
6 benchmarks · 1 papers
First Incorrect Step Identification
6 benchmarks · 3 papers
Zero-shot Prediction
6 benchmarks · 1 papers
Multi-turn Visual Question Answering
6 benchmarks · 10 papers
Multistep Reasoning
6 benchmarks · 6 papers
Knowledge-Grounded Dialogue Generation
6 benchmarks · 2 papers
Deep Research Evaluation
6 benchmarks · 2 papers
Quality Evaluation
6 benchmarks · 1 papers
Episodic Memory Recall
6 benchmarks · 12 papers
General Instruction Following
6 benchmarks · 4 papers
Instruction Generation
Page 13 of 154
Previous
Next