Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 43 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
2 benchmarks · 1 papers
Task Completion and Utility
2 benchmarks · 2 papers
Role-playing agent evaluation
2 benchmarks · 1 papers
Abstention Classification
2 benchmarks · 1 papers
Attribute-Controlled Dialogue Generation
2 benchmarks · 1 papers
Multi-step reasoning and knowledge retrieval
2 benchmarks · 1 papers
Proactive information probing
2 benchmarks · 2 papers
Length Extrapolation
2 benchmarks · 1 papers
Price Negotiation
2 benchmarks · 2 papers
Rationale alignment
2 benchmarks · 2 papers
Generative tasks
2 benchmarks · 1 papers
Arabic Cultural Value Alignment
2 benchmarks · 2 papers
Dense prompt following
2 benchmarks · 1 papers
Psychology counseling response generation
2 benchmarks · 1 papers
Reasoning Fidelity
2 benchmarks · 2 papers
Social Intelligence Reasoning
2 benchmarks · 1 papers
Corrective Instruction Generation
2 benchmarks · 1 papers
Multi-turn Jailbreak Evaluation
2 benchmarks · 1 papers
Skill Invocation Safety Auditing
2 benchmarks · 1 papers
Real-trace replay validation
2 benchmarks · 2 papers
Grade-school reasoning
2 benchmarks · 1 papers
Math Word Problem Preference
2 benchmarks · 1 papers
Broad Knowledge Question Answering
2 benchmarks · 1 papers
Logical Reasoning and Reading Comprehension
2 benchmarks · 1 papers
Loop Suppression
2 benchmarks · 4 papers
Response Harmfulness Classification
2 benchmarks · 2 papers
Aggregate Reasoning Evaluation
2 benchmarks · 2 papers
User Preference
2 benchmarks · 1 papers
Context Traceback
2 benchmarks · 2 papers
Long-context Multi-modal Understanding
2 benchmarks · 1 papers
Rule Following
2 benchmarks · 1 papers
Language-conditioned simulation
2 benchmarks · 1 papers
Consultation Capability Evaluation
Page 43 of 154
Previous
Next