Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 108 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
1 benchmarks · 1 papers
Chain-of-Thought Compression
1 benchmarks · 1 papers
Knowledge boundary assessment
1 benchmarks · 1 papers
Science Q&A
1 benchmarks · 1 papers
Overall Critique and Refinement
1 benchmarks · 1 papers
Co-Creativity Assessment
1 benchmarks · 1 papers
Incident Diagnosis and Resolution
1 benchmarks · 1 papers
Chit-chat Dialogue Generation
1 benchmarks · 1 papers
Egregious unfaithfulness
1 benchmarks · 1 papers
Multi-task Success Evaluation
1 benchmarks · 1 papers
Action-conditioned reasoning
1 benchmarks · 1 papers
Multimodal Evaluation (Cognition/Summary)
1 benchmarks · 1 papers
Value Fusion
1 benchmarks · 1 papers
Instruction Following Safety
1 benchmarks · 1 papers
Pairwise Safety Evaluation
1 benchmarks · 1 papers
Long-horizon memory-based reasoning
1 benchmarks · 1 papers
Instruction Hierarchy
1 benchmarks · 1 papers
Chain Generation
1 benchmarks · 1 papers
Tool-using Reasoning
1 benchmarks · 1 papers
Multi-task Language Understanding and Reasoning
1 benchmarks · 1 papers
Discrimination between Good Faith and Problematic agents (Peer Review)
1 benchmarks · 1 papers
Explanation Alignment
1 benchmarks · 1 papers
Diagnostic Dialogue
1 benchmarks · 1 papers
Reasoning and Generative Tasks
1 benchmarks · 1 papers
Length-controlled text generation
1 benchmarks · 1 papers
Intent Coverage and Efficiency
1 benchmarks · 1 papers
Best Research Idea Selection
1 benchmarks · 1 papers
Multi-task multimodal understanding
1 benchmarks · 1 papers
LLM-to-LLM Persuasion
1 benchmarks · 1 papers
Consistent Response Score
1 benchmarks · 1 papers
Reasoning over conflicting evidence
1 benchmarks · 1 papers
Practicality Assessment
1 benchmarks · 1 papers
Sycophancy Correction Receptiveness
Page 108 of 154
Previous
Next