Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 52 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
2 benchmarks · 2 papers
Proof writing
2 benchmarks · 5 papers
Role-Play Evaluation
2 benchmarks · 3 papers
Role-Play
2 benchmarks · 1 papers
Reward Model Controllability
2 benchmarks · 1 papers
Generalization to Unseen Preferences
2 benchmarks · 1 papers
Rational Ordering Evaluation
2 benchmarks · 1 papers
Ordinal Preference Alignment
2 benchmarks · 2 papers
Coherence evaluation
2 benchmarks · 1 papers
Reward-wise QA fairness and alignment
2 benchmarks · 1 papers
Expert Preference Pairwise
2 benchmarks · 1 papers
Robot Failure Analysis (MCQ)
2 benchmarks · 1 papers
Scientific Verification
2 benchmarks · 1 papers
Chit-chat conversation evaluation correlation
2 benchmarks · 1 papers
Multi-hop Retrieval-Augmented Generation
2 benchmarks · 1 papers
Reference-free Conversation Evaluation
2 benchmarks · 3 papers
Next Step Generation
2 benchmarks · 2 papers
Language Understanding and Question Answering
2 benchmarks · 1 papers
LLM steering evaluation
2 benchmarks · 2 papers
General Reasoning Evaluation
2 benchmarks · 1 papers
Policy Alignment
2 benchmarks · 1 papers
Thought Generation
2 benchmarks · 1 papers
Mathematical and Knowledge Reasoning
2 benchmarks · 3 papers
Math & Knowledge
2 benchmarks · 1 papers
Friend Recommendation
2 benchmarks · 1 papers
Dialogue Annotation
2 benchmarks · 2 papers
Multi-Agent Collaboration
2 benchmarks · 2 papers
Model Evaluation
2 benchmarks · 2 papers
Response Evaluation
2 benchmarks · 1 papers
Accuracy Evaluation
2 benchmarks · 2 papers
Verifiable Instruction Following
2 benchmarks · 1 papers
Adversarial Toxicity Refusal
2 benchmarks · 1 papers
Open-domain Dialogue Evaluation
Page 52 of 154
Previous
Next