Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 64 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
2 benchmarks · 2 papers
Arithmetic Planning
2 benchmarks · 1 papers
Math Question Verification
2 benchmarks · 1 papers
CBT Conversation Generation
2 benchmarks · 2 papers
Multi-agent Negotiation
2 benchmarks · 3 papers
Complex Multi-step Reasoning
2 benchmarks · 3 papers
Multi-domain Knowledge and Reasoning
2 benchmarks · 1 papers
Exam
2 benchmarks · 2 papers
MultiModal Long-Context Understanding
2 benchmarks · 1 papers
Autoformalization and Proving
2 benchmarks · 1 papers
Alignment Reward Evaluation
2 benchmarks · 1 papers
Scientific Review Feedback Generation
2 benchmarks · 2 papers
Reward Scoring
2 benchmarks · 1 papers
Prefill-stage hallucination risk detection
2 benchmarks · 1 papers
Prompt Hygiene Evaluation
2 benchmarks · 1 papers
Long-term preference alignment
2 benchmarks · 1 papers
High-level instruction execution
2 benchmarks · 1 papers
Multilingual Reward Modeling
2 benchmarks · 1 papers
Multi-modal Dialogue
2 benchmarks · 2 papers
Helpfulness Assessment
2 benchmarks · 2 papers
Video Preference Evaluation
2 benchmarks · 1 papers
Chat Performance
2 benchmarks · 1 papers
Large Language Model Debiasing
2 benchmarks · 2 papers
LLM Agent Task Completion
2 benchmarks · 1 papers
Conflict Measurement
2 benchmarks · 2 papers
Rationale Faithfulness Evaluation
2 benchmarks · 4 papers
Multi-discipline Understanding
2 benchmarks · 2 papers
Context Management
2 benchmarks · 2 papers
Audio Instruction Following
2 benchmarks · 1 papers
Shot-Language Understanding
2 benchmarks · 2 papers
Web-based Reasoning
2 benchmarks · 1 papers
Social Norm Dialogue Generation
2 benchmarks · 2 papers
Zero-shot Reasoning and Knowledge
Page 64 of 154
Previous
Next