Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 132 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
1 benchmarks · 1 papers
General Description
1 benchmarks · 1 papers
Diversity of solution strategies
1 benchmarks · 1 papers
Alignment with Human Preferences
1 benchmarks · 1 papers
Multi-task model alignment and mixing
1 benchmarks · 2 papers
Refusal Control
1 benchmarks · 1 papers
Reasoning Correction
1 benchmarks · 1 papers
Persona Control
1 benchmarks · 1 papers
General Intelligence
1 benchmarks · 1 papers
Contextual
1 benchmarks · 1 papers
Long-context Instruction Following
1 benchmarks · 1 papers
Intent Mismatch Detection
1 benchmarks · 1 papers
Expert-Level Human Knowledge Reasoning
1 benchmarks · 1 papers
Deep Search and Research Reasoning
1 benchmarks · 1 papers
Multi-step Reasoning and Factuality
1 benchmarks · 1 papers
Online Troubleshooting
1 benchmarks · 1 papers
Correlation analysis with human preferences
1 benchmarks · 1 papers
Multi-task long-context understanding
1 benchmarks · 1 papers
Long-horizon memory reasoning and retrieval
1 benchmarks · 1 papers
Atomic Memory Recall
1 benchmarks · 1 papers
Reasoning-chain quality evaluation
1 benchmarks · 1 papers
Recovery Evaluation
1 benchmarks · 1 papers
Complex Factual Reasoning
1 benchmarks · 1 papers
Instruction Dataset Quality Evaluation
1 benchmarks · 1 papers
Blind pairwise comparison
1 benchmarks · 1 papers
Logic-heavy Reasoning
1 benchmarks · 2 papers
LLM Instruction Tuning
1 benchmarks · 1 papers
Role-Playing Ability
1 benchmarks · 1 papers
Multi-Turn Consistency
1 benchmarks · 1 papers
Theory of Mind Inference
1 benchmarks · 1 papers
Prompt Robustness Evaluation
1 benchmarks · 1 papers
Role-Play Faithfulness
1 benchmarks · 1 papers
Memory-Augmented GUI Interaction
Page 132 of 154
Previous
Next