Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 93 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
1 benchmarks · 1 papers
VLM Pairwise Preference
1 benchmarks · 3 papers
Grade School Math Word Problem Solving
1 benchmarks · 1 papers
Contextual Knowledge Adaptation (Prior-Conflicting)
1 benchmarks · 1 papers
Slide Deck Generation Quality Evaluation
1 benchmarks · 1 papers
LLM Safety
1 benchmarks · 1 papers
Tool-use Execution
1 benchmarks · 1 papers
Expert Knowledge Evaluation
1 benchmarks · 1 papers
Planning and Tool Use
1 benchmarks · 1 papers
Output Specificity
1 benchmarks · 1 papers
Multi-party Dialogue Question Answering
1 benchmarks · 1 papers
Single-Hop Reasoning
1 benchmarks · 1 papers
Multi-domain Question Answering
1 benchmarks · 2 papers
Overall Reasoning (Average)
1 benchmarks · 1 papers
Factual Answering
1 benchmarks · 1 papers
Multi-Token Prediction
1 benchmarks · 1 papers
Emotion Steering
1 benchmarks · 1 papers
Validity Certification
1 benchmarks · 1 papers
Multi-choice Evaluation
1 benchmarks · 1 papers
Relative Robustness Analysis
1 benchmarks · 1 papers
Cross-size lineage similarity detection
1 benchmarks · 1 papers
Cognitive Task Assessment
1 benchmarks · 1 papers
Hidden Intent Inference
1 benchmarks · 1 papers
Agentic Person Search (Spatial Reasoning)
1 benchmarks · 1 papers
Agentic Person Search (Temporal Reasoning)
1 benchmarks · 1 papers
Knowledge-intensive language tasks evaluation
1 benchmarks · 1 papers
Shuffle Dyck
1 benchmarks · 1 papers
Pairwise evaluator agreement with human judgment
1 benchmarks · 1 papers
Simulated Patient Portrayal
1 benchmarks · 1 papers
Reward-Guided Generation
1 benchmarks · 1 papers
End-to-end dialogue generation
1 benchmarks · 1 papers
Reasoned Question Answering
1 benchmarks · 1 papers
Interaction integrity monitoring
Page 93 of 154
Previous
Next