Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 72 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
1 benchmarks · 1 papers
Hallucination Detection (Math Word Problems)
1 benchmarks · 1 papers
Hallucination Detection (Self-contradictory Hallucinations)
1 benchmarks · 1 papers
Language-conditioned long-horizon robot manipulation
1 benchmarks · 1 papers
Open-ended Visual Chat
1 benchmarks · 1 papers
Conversational Agent Interaction
1 benchmarks · 1 papers
Length Following
1 benchmarks · 1 papers
Long-generation
1 benchmarks · 1 papers
Short-length Instruction Following
1 benchmarks · 1 papers
Psychotherapy Dialogue Evaluation
1 benchmarks · 1 papers
Clinical Medical Reasoning Evaluation
1 benchmarks · 1 papers
Zero-shot Task Completion
1 benchmarks · 1 papers
Language Instruction Following
1 benchmarks · 1 papers
Conversation2Note
1 benchmarks · 2 papers
Overall Language Performance
1 benchmarks · 1 papers
Zero-shot long-context reasoning
1 benchmarks · 1 papers
Synthetic Dialogue Generation Evaluation
1 benchmarks · 2 papers
Factual Accuracy and Reasoning
1 benchmarks · 1 papers
Long-context reasoning (Pairs)
1 benchmarks · 1 papers
Client Simulation Receptivity Consistency
1 benchmarks · 1 papers
Overall reasoning performance
1 benchmarks · 1 papers
General Language Model Capability
1 benchmarks · 1 papers
RLHF Alignment Evaluation
1 benchmarks · 1 papers
High-level instruction following
1 benchmarks · 1 papers
College-level Multimodal Understanding
1 benchmarks · 1 papers
Emotional Support Conversation Seeker Simulation
1 benchmarks · 1 papers
Multi-agent aggregation
1 benchmarks · 1 papers
Dialogue Diversity Evaluation
1 benchmarks · 1 papers
Grounded Chess Reasoning
1 benchmarks · 1 papers
Longform generation of biographies
1 benchmarks · 1 papers
Logical reasoning multi-choice QA
1 benchmarks · 1 papers
Seeker Utterance Generation
1 benchmarks · 1 papers
Mobile Use
Page 72 of 154
Previous
Next