Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 62 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
2 benchmarks · 2 papers
Multiple-choice commonsense reasoning
2 benchmarks · 1 papers
Summary Factuality Evaluation
2 benchmarks · 2 papers
Social interaction simulation
2 benchmarks · 2 papers
Arithmetic Planning
2 benchmarks · 3 papers
Complex Multi-step Reasoning
2 benchmarks · 2 papers
MultiModal Long-Context Understanding
2 benchmarks · 2 papers
Reward Scoring
2 benchmarks · 1 papers
Alignment Reward Evaluation
2 benchmarks · 1 papers
High-level instruction execution
2 benchmarks · 1 papers
Autoformalization and Proving
2 benchmarks · 2 papers
General Language Intelligence
2 benchmarks · 2 papers
Grade-school mathematical reasoning
2 benchmarks · 2 papers
Helpfulness Assessment
2 benchmarks · 1 papers
Dialogue Aspect-based Sentiment Quadruple Extraction
2 benchmarks · 1 papers
Prompt Hygiene Evaluation
2 benchmarks · 1 papers
Generative Multiple-choice Question Answering
2 benchmarks · 1 papers
Hallucination Tracing
2 benchmarks · 1 papers
Generative multiple-choice
2 benchmarks · 1 papers
Long-Horizon Search Intelligence
2 benchmarks · 1 papers
Evaluation Reliability
2 benchmarks · 1 papers
Prefill-stage hallucination risk detection
2 benchmarks · 1 papers
Long-form reasoning
2 benchmarks · 2 papers
Persona Consistency
2 benchmarks · 2 papers
Audio Instruction Following
2 benchmarks · 1 papers
Conflict Measurement
2 benchmarks · 1 papers
Behavior Generation
2 benchmarks · 2 papers
Rationale Faithfulness Evaluation
2 benchmarks · 2 papers
Confidence Alignment
2 benchmarks · 2 papers
Context Management
2 benchmarks · 1 papers
Instruction State Tracking
2 benchmarks · 1 papers
Shot-Language Understanding
2 benchmarks · 2 papers
Counseling Dialogue Evaluation
Page 62 of 154
Previous
Next