Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 9 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
9 benchmarks · 7 papers
Overrefusal evaluation
9 benchmarks · 8 papers
Multilingual Reasoning
9 benchmarks · 8 papers
Multiple-Choice Classification
9 benchmarks · 1 papers
Functionally Diverse Response Generation
9 benchmarks · 3 papers
Win Rate Evaluation
9 benchmarks · 2 papers
LLM Red-teaming
9 benchmarks · 1 papers
Comprehensiveness
9 benchmarks · 4 papers
Emotional Support Dialogue
9 benchmarks · 6 papers
Scientific Idea Generation
9 benchmarks · 2 papers
Dialog Response Generation
8 benchmarks · 1 papers
Task-Focused Dialogue
8 benchmarks · 2 papers
Emotional Reasoning
8 benchmarks · 4 papers
Constraint Following
8 benchmarks · 4 papers
Sycophancy Detection
8 benchmarks · 2 papers
Output Length Prediction
8 benchmarks · 2 papers
Tool-use instruction following
8 benchmarks · 12 papers
General Knowledge Question Answering
8 benchmarks · 14 papers
Multimodal Conversation
8 benchmarks · 1 papers
Multi-path speculative decoding
8 benchmarks · 10 papers
Language Modeling Evaluation
8 benchmarks · 3 papers
Conversational Memory Retrieval
8 benchmarks · 4 papers
Input Moderation
8 benchmarks · 1 papers
Global resource management and long-term strategic planning
8 benchmarks · 25 papers
Long-term memory evaluation
8 benchmarks · 3 papers
System Prompt Extraction
8 benchmarks · 1 papers
Multi-instruction Image Editing
8 benchmarks · 8 papers
General LLM Evaluation
8 benchmarks · 9 papers
General Language Evaluation
8 benchmarks · 3 papers
Lifelong Knowledge Editing
8 benchmarks · 3 papers
Memorization mitigation
8 benchmarks · 6 papers
LLM Safety Evaluation
8 benchmarks · 1 papers
Problem Solving and Unsolvability Detection
Page 9 of 154
Previous
Next