Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 39 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
3 benchmarks · 5 papers
Multistep Soft Reasoning
3 benchmarks · 3 papers
Meeting Planning
3 benchmarks · 1 papers
Multi-agent Cognitive Orchestration
3 benchmarks · 4 papers
Conversational Memory
3 benchmarks · 1 papers
Reasoning failure prediction
3 benchmarks · 1 papers
Preference Labeling
3 benchmarks · 3 papers
Prefill
3 benchmarks · 2 papers
Continual Instruction Following
3 benchmarks · 4 papers
Out-of-Distribution Reasoning
3 benchmarks · 1 papers
Sycophancy Assessment
3 benchmarks · 1 papers
Language-Conditioned Tasks
3 benchmarks · 1 papers
Reasoning trace quality evaluation
3 benchmarks · 1 papers
Bayesian Assessment of Sycophancy
3 benchmarks · 1 papers
Linear Concept Accessibility and Steering
3 benchmarks · 4 papers
Multi-Turn Function Calling
3 benchmarks · 3 papers
Instruction-based Video Editing
3 benchmarks · 1 papers
Tool-use Inference
3 benchmarks · 1 papers
Fine-grained Knowledge Recall
3 benchmarks · 1 papers
Multi-turn dialogue routing
3 benchmarks · 3 papers
Task Decomposition
3 benchmarks · 1 papers
Goal-relevance Evaluation
3 benchmarks · 1 papers
Rubric satisfaction evaluation
3 benchmarks · 1 papers
Skill Learning
3 benchmarks · 3 papers
Safety-Utility Trade-off Evaluation
3 benchmarks · 2 papers
Social Deduction Game Gameplay
3 benchmarks · 1 papers
Long-form factuality evaluation
3 benchmarks · 1 papers
LoRA Adapter Transfer
3 benchmarks · 2 papers
Knowledge Injection
3 benchmarks · 1 papers
Single-Turn Mathematical Reasoning
3 benchmarks · 2 papers
Multi-turn Jailbreak
3 benchmarks · 1 papers
Self-confidence estimation
3 benchmarks · 4 papers
Multimodal Medical Reasoning
Page 39 of 154
Previous
Next