Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 5 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
15 benchmarks · 28 papers
General Knowledge Evaluation
15 benchmarks · 4 papers
Arithmetic Addition
15 benchmarks · 13 papers
Helpfulness evaluation
14 benchmarks · 2 papers
Short-form generation
14 benchmarks · 5 papers
Faithfulness Detection
14 benchmarks · 9 papers
Long Context
14 benchmarks · 16 papers
General Capability
14 benchmarks · 18 papers
Long-context Memory Evaluation
14 benchmarks · 3 papers
Controllability
14 benchmarks · 8 papers
Safety and Utility Evaluation
14 benchmarks · 20 papers
Chat
14 benchmarks · 6 papers
Cultural Alignment
14 benchmarks · 7 papers
Response Classification
14 benchmarks · 2 papers
Graph Classification Explanation
14 benchmarks · 10 papers
Multimodal Reward Modeling
14 benchmarks · 9 papers
Continual Instruction Tuning
14 benchmarks · 3 papers
Domain Reasoning
14 benchmarks · 4 papers
Long-form QA
14 benchmarks · 7 papers
Idea Generation
14 benchmarks · 12 papers
Factual Question Answering
13 benchmarks · 11 papers
Mathematics Reasoning
13 benchmarks · 4 papers
Misalignment Detection
13 benchmarks · 3 papers
Long-context Memory Retrieval and Reasoning
13 benchmarks · 10 papers
Interactive Video Generation
13 benchmarks · 3 papers
Story Continuation
13 benchmarks · 10 papers
LLM Inference Efficiency
13 benchmarks · 2 papers
Model Ranking Prediction
13 benchmarks · 4 papers
Proactive Assistance
13 benchmarks · 14 papers
Downstream Task Evaluation
13 benchmarks · 12 papers
Zero-shot Question Answering
13 benchmarks · 22 papers
Truthful Question Answering
13 benchmarks · 5 papers
Dialog
Page 5 of 154
Previous
Next