Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 53 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
2 benchmarks · 1 papers
Open-domain Dialogue Evaluation
2 benchmarks · 3 papers
Long-context language evaluation
2 benchmarks · 1 papers
Reference-free Conversation Evaluation
2 benchmarks · 1 papers
Expert Preference Pairwise
2 benchmarks · 2 papers
Consistency Analysis
2 benchmarks · 1 papers
Multi-hop Retrieval-Augmented Generation
2 benchmarks · 1 papers
Chit-chat conversation evaluation correlation
2 benchmarks · 2 papers
Prompt Selection
2 benchmarks · 1 papers
Pun Explanation
2 benchmarks · 1 papers
Generalization to Unseen Preferences
2 benchmarks · 1 papers
Ordinal Preference Alignment
2 benchmarks · 1 papers
Friend Recommendation
2 benchmarks · 1 papers
Dialogue Annotation
2 benchmarks · 1 papers
LLM steering evaluation
2 benchmarks · 2 papers
Multi-Agent Collaboration
2 benchmarks · 2 papers
Game Solving
2 benchmarks · 1 papers
Adversarial Toxicity Refusal
2 benchmarks · 2 papers
Language Understanding and Question Answering
2 benchmarks · 1 papers
Hallucination-oriented Video Understanding
2 benchmarks · 1 papers
Question Answering with Clarification
2 benchmarks · 1 papers
Crisis response generation
2 benchmarks · 3 papers
Mathematical Reasoning Process Evaluation
2 benchmarks · 1 papers
Agent Planning and API Calling
2 benchmarks · 1 papers
Correctness Assessment
2 benchmarks · 1 papers
Offline Constrained RLHF
2 benchmarks · 1 papers
Static Multi-Session QA
2 benchmarks · 1 papers
Short-form open-domain QA
2 benchmarks · 1 papers
Reasoning and Knowledge Assessment
2 benchmarks · 2 papers
Long-form Retrieval-Augmented Generation
2 benchmarks · 2 papers
Long-context Multi-modal Understanding
2 benchmarks · 1 papers
Personality Recovery
2 benchmarks · 1 papers
Instruction-Guided Grading
Page 53 of 154
Previous
Next