Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 3 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
30 benchmarks · 2 papers
Detection of LLM generated text
30 benchmarks · 14 papers
Personalization
29 benchmarks · 19 papers
Instruction Following Evaluation
28 benchmarks · 11 papers
Personalized Text Generation
28 benchmarks · 2 papers
Selective Generation
27 benchmarks · 19 papers
Preference Evaluation
27 benchmarks · 20 papers
Human Preference Alignment
27 benchmarks · 11 papers
Safety Alignment Evaluation
27 benchmarks · 19 papers
Knowledge Retention
26 benchmarks · 18 papers
Multi-step Reasoning
26 benchmarks · 14 papers
Sycophancy Evaluation
25 benchmarks · 15 papers
LLM Unlearning
25 benchmarks · 9 papers
Negotiation
25 benchmarks · 13 papers
Long-form generation
25 benchmarks · 22 papers
Geometric Reasoning
24 benchmarks · 21 papers
Dialogue
24 benchmarks · 31 papers
Multilingual Mathematical Reasoning
23 benchmarks · 7 papers
Prompt Classification
23 benchmarks · 7 papers
Malicious Prompt Detection
23 benchmarks · 15 papers
Generative Question Answering
23 benchmarks · 1 papers
Cultural safety evaluation
22 benchmarks · 11 papers
Role-playing
22 benchmarks · 19 papers
Long-context retrieval and reasoning
22 benchmarks · 18 papers
Knowledge-intensive reasoning
22 benchmarks · 16 papers
Hallucination Mitigation
22 benchmarks · 30 papers
General AI Assistant Tasks
21 benchmarks · 3 papers
LLM-generated content detection
21 benchmarks · 18 papers
Over-refusal evaluation
21 benchmarks · 11 papers
World Knowledge
20 benchmarks · 34 papers
Truthfulness
20 benchmarks · 1 papers
Insight-level Evaluation
20 benchmarks · 9 papers
Fine-tuning
Page 3 of 154
Previous
Next