Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 14 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
6 benchmarks · 1 papers
LLM Watermark Spoofing
6 benchmarks · 4 papers
Long-context Summarization
6 benchmarks · 2 papers
Motivational Interviewing Dialogue Evaluation
6 benchmarks · 3 papers
Personalized Dialogue Generation
6 benchmarks · 7 papers
Multi-Task Reasoning
6 benchmarks · 12 papers
General Instruction Following
6 benchmarks · 1 papers
First Incorrect Step Identification
6 benchmarks · 1 papers
Persona Bias Evaluation
6 benchmarks · 2 papers
Generative Performance
6 benchmarks · 4 papers
Presentation Generation
6 benchmarks · 10 papers
Multistep Reasoning
6 benchmarks · 5 papers
User Interruption
6 benchmarks · 7 papers
Multi-task Language Evaluation
6 benchmarks · 3 papers
Model Alignment
6 benchmarks · 4 papers
Generalization Reasoning
6 benchmarks · 1 papers
Textual Quality Evaluation
6 benchmarks · 4 papers
General Assistant Tasks
6 benchmarks · 3 papers
LLM Prefill Throughput
6 benchmarks · 3 papers
Zero-shot Prediction
6 benchmarks · 3 papers
Speech Instruction-Following
6 benchmarks · 2 papers
Robot Instruction Following
6 benchmarks · 2 papers
Quality Evaluation
6 benchmarks · 2 papers
Knowledge-Grounded Conversation
6 benchmarks · 1 papers
Episodic Memory Recall
6 benchmarks · 1 papers
Episodic Memory Retrieval
6 benchmarks · 3 papers
Countdown
6 benchmarks · 1 papers
Failure Detection and Reasoning
6 benchmarks · 4 papers
Refusal Detection
6 benchmarks · 4 papers
Multi-turn Dialogue Generation
6 benchmarks · 3 papers
Question Answering with Abstention
6 benchmarks · 1 papers
Long-horizon agentic tasks
6 benchmarks · 1 papers
Multi-turn Visual Question Answering
Page 14 of 154
Previous
Next