Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 124 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
1 benchmarks · 1 papers
Chess Reasoning Quality Evaluation
1 benchmarks · 1 papers
Hallucination / grounding
1 benchmarks · 2 papers
General reasoning / multi-discipline
1 benchmarks · 1 papers
Factual Grounding
1 benchmarks · 1 papers
Commonsense Reasoning and Knowledge Question Answering
1 benchmarks · 1 papers
Response Generation Quality
1 benchmarks · 1 papers
Conversational Alignment
1 benchmarks · 1 papers
Comparative Ranking
1 benchmarks · 1 papers
Intent Reflection
1 benchmarks · 1 papers
Subjective evaluation of agent feedback
1 benchmarks · 1 papers
Sponsored product promotion
1 benchmarks · 1 papers
Long-term Question Answering
1 benchmarks · 1 papers
Adversarial Hallucination Attack
1 benchmarks · 1 papers
Open-ended MCQA
1 benchmarks · 1 papers
Academic Examination
1 benchmarks · 1 papers
Constraint Satisfaction Plan Generation
1 benchmarks · 1 papers
Closed-ended Task Evaluation
1 benchmarks · 1 papers
General-purpose Behavior
1 benchmarks · 1 papers
Intermediate answer generation
1 benchmarks · 1 papers
Web Reasoning
1 benchmarks · 1 papers
Model Merging Evaluation
1 benchmarks · 1 papers
Multi-task and Multi-image Reasoning
1 benchmarks · 1 papers
Overall Language Model Evaluation
1 benchmarks · 1 papers
Prompt Enhancement for Video Generation
1 benchmarks · 1 papers
Skill Knowledge
1 benchmarks · 1 papers
Prompt Enhancement for Text-to-Video Generation
1 benchmarks · 1 papers
Distractor Selection
1 benchmarks · 1 papers
Abstention Trigger
1 benchmarks · 1 papers
Interactive Text Editing (Continuation)
1 benchmarks · 2 papers
Joke generation
1 benchmarks · 1 papers
Post-training Evaluation
1 benchmarks · 1 papers
Intent Shifting
Page 124 of 154
Previous
Next