Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 8 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
10 benchmarks · 2 papers
Multimodal Preference Evaluation
10 benchmarks · 1 papers
MCQ Classification
10 benchmarks · 1 papers
Harmful Request Compliance
10 benchmarks · 1 papers
Downstream accuracy extrapolation
10 benchmarks · 1 papers
Scientific Foundation Model Collaboration
10 benchmarks · 1 papers
Instruction-following clustering
10 benchmarks · 4 papers
Decoding Throughput
10 benchmarks · 2 papers
Diversity
9 benchmarks · 2 papers
LLM Red-teaming
9 benchmarks · 1 papers
Functionally Diverse Response Generation
9 benchmarks · 9 papers
General Capability Evaluation
9 benchmarks · 9 papers
Jailbreak Safety Evaluation
9 benchmarks · 4 papers
Patient Simulation
9 benchmarks · 9 papers
Multiple-Choice Reasoning
9 benchmarks · 9 papers
Response Quality Evaluation
9 benchmarks · 3 papers
Win Rate Evaluation
9 benchmarks · 8 papers
Long-term Memory Question Answering
9 benchmarks · 8 papers
General Language Modeling
9 benchmarks · 4 papers
Multi-Query Associative Recall
9 benchmarks · 5 papers
Long-form writing
9 benchmarks · 17 papers
Writing
9 benchmarks · 1 papers
Model Ranking
9 benchmarks · 12 papers
Zero-shot Commonsense Reasoning
9 benchmarks · 8 papers
Emotional Intelligence Evaluation
9 benchmarks · 17 papers
Multi-discipline Reasoning
9 benchmarks · 15 papers
Multi-turn Instruction Following
9 benchmarks · 5 papers
Open-domain dialogue
9 benchmarks · 7 papers
Task 2
9 benchmarks · 5 papers
Conflict Resolution
9 benchmarks · 4 papers
Scientific ideation
9 benchmarks · 8 papers
Multilingual Reasoning
9 benchmarks · 8 papers
Multiple-Choice Classification
Page 8 of 154
Previous
Next