Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 63 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
2 benchmarks · 2 papers
General Downstream Evaluation
2 benchmarks · 2 papers
MultiModal Long-Context Understanding
2 benchmarks · 1 papers
Alignment Reward Evaluation
2 benchmarks · 1 papers
Personalized Review Writing
2 benchmarks · 2 papers
LLM-judge evaluation
2 benchmarks · 1 papers
Evaluation Reliability
2 benchmarks · 1 papers
Autoformalization and Proving
2 benchmarks · 1 papers
Prefill-stage hallucination risk detection
2 benchmarks · 1 papers
Open-domain dialogue red teaming
2 benchmarks · 2 papers
Helpfulness Assessment
2 benchmarks · 1 papers
Prompt continuation
2 benchmarks · 1 papers
LLM Agent Reasoning
2 benchmarks · 2 papers
Aggregated LLM Evaluation
2 benchmarks · 2 papers
Multiple-choice commonsense reasoning
2 benchmarks · 1 papers
Prompt Hygiene Evaluation
2 benchmarks · 1 papers
Self-introduction generation
2 benchmarks · 2 papers
Dialogue Quality
2 benchmarks · 1 papers
Large Language Model Debiasing
2 benchmarks · 1 papers
Persona Role-Playing Faithfulness
2 benchmarks · 2 papers
Rationale Faithfulness Evaluation
2 benchmarks · 1 papers
Conflict Measurement
2 benchmarks · 2 papers
Audio Instruction Following
2 benchmarks · 1 papers
Math Question Verification
2 benchmarks · 2 papers
Context Management
2 benchmarks · 1 papers
CBT Conversation Generation
2 benchmarks · 1 papers
Shot-Language Understanding
2 benchmarks · 2 papers
Zero-shot Reasoning and Knowledge
2 benchmarks · 3 papers
Multi-domain Knowledge and Reasoning
2 benchmarks · 1 papers
Exam
2 benchmarks · 1 papers
Average across tasks
2 benchmarks · 1 papers
Scientific Review Feedback Generation
2 benchmarks · 2 papers
Mathematical Calculation
Page 63 of 154
Previous
Next