Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 24 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
4 benchmarks · 2 papers
Feasibility classification
4 benchmarks · 1 papers
Misaligned Task Learning
4 benchmarks · 1 papers
Reasoning-Aware Threat Detection
4 benchmarks · 1 papers
LLM-as-judge head-to-head comparison
4 benchmarks · 2 papers
Personality Understanding
4 benchmarks · 4 papers
Comprehensive Examination
4 benchmarks · 1 papers
Tool-use behavioral profiling
4 benchmarks · 2 papers
Chat & Instruction Following
4 benchmarks · 1 papers
Transitive reasoning
4 benchmarks · 1 papers
Multilingual Reasoning and Language Understanding
4 benchmarks · 1 papers
Input-based feature description evaluation
4 benchmarks · 1 papers
Output-based feature description evaluation
4 benchmarks · 1 papers
Function Vector Evaluation
4 benchmarks · 2 papers
Dialogue Agent Interaction
4 benchmarks · 2 papers
Long-Horizon Stability
4 benchmarks · 4 papers
Utility assessment
4 benchmarks · 1 papers
Explanation self-consistency
4 benchmarks · 2 papers
Pairwise Evaluation
4 benchmarks · 1 papers
Deductive logical reasoning
4 benchmarks · 4 papers
Output Diversity
4 benchmarks · 1 papers
Therapeutic Game Evaluation
4 benchmarks · 2 papers
Variable Tracking
4 benchmarks · 1 papers
Disruption
4 benchmarks · 1 papers
CoT Soundness Evaluation
4 benchmarks · 1 papers
Long-context Processing
4 benchmarks · 2 papers
LLM Personalization
4 benchmarks · 5 papers
General Intelligence Evaluation
4 benchmarks · 1 papers
Multi-party Dialogue Generation
4 benchmarks · 1 papers
Long-Context Stability & Retrieval
4 benchmarks · 1 papers
CoT Naturalness
4 benchmarks · 1 papers
Information Analysis and Processing
4 benchmarks · 3 papers
Memory Capability Evaluation
Page 24 of 154
Previous
Next