Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 51 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
2 benchmarks · 1 papers
Citation Coverage Evaluation
2 benchmarks · 1 papers
Open-ended QA Response Ranking
2 benchmarks · 1 papers
Reward Model Controllability
2 benchmarks · 1 papers
Generalization to Unseen Preferences
2 benchmarks · 2 papers
Proof writing
2 benchmarks · 1 papers
Preference Profile Estimation
2 benchmarks · 2 papers
General Reasoning Average
2 benchmarks · 1 papers
LLM Inference Throughput
2 benchmarks · 1 papers
Reward-wise QA fairness and alignment
2 benchmarks · 1 papers
Expert Preference Pairwise
2 benchmarks · 1 papers
Reference-free Conversation Evaluation
2 benchmarks · 1 papers
Ordinal Preference Alignment
2 benchmarks · 1 papers
Diversity measurement correlation
2 benchmarks · 1 papers
Chat Fine-tuning
2 benchmarks · 1 papers
Multi-hop Retrieval-Augmented Generation
2 benchmarks · 1 papers
Memorization Reduction
2 benchmarks · 3 papers
Synthetic Tasks
2 benchmarks · 1 papers
Robustness against harmful content generation
2 benchmarks · 1 papers
Chit-chat conversation evaluation correlation
2 benchmarks · 1 papers
Friend Recommendation
2 benchmarks · 1 papers
Math Word Problem Reasoning
2 benchmarks · 1 papers
Long-generation reasoning
2 benchmarks · 4 papers
Long-context performance evaluation
2 benchmarks · 2 papers
General Capability Retention
2 benchmarks · 2 papers
Hallucination Annotation
2 benchmarks · 2 papers
Multilingual Commonsense Reasoning
2 benchmarks · 1 papers
Dialogue Annotation
2 benchmarks · 2 papers
Multi-Agent Collaboration
2 benchmarks · 1 papers
Human Preferences
2 benchmarks · 1 papers
LLM steering evaluation
2 benchmarks · 5 papers
Role-Play Evaluation
2 benchmarks · 3 papers
Role-Play
Page 51 of 154
Previous
Next