Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 59 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
2 benchmarks · 1 papers
Sequential Multi-Task Learning
2 benchmarks · 1 papers
Interactive query answering
2 benchmarks · 1 papers
User Perceived Understandability
2 benchmarks · 1 papers
General Knowledge and Instruction Following
2 benchmarks · 2 papers
Secure LLM Agent Task Completion
2 benchmarks · 2 papers
General AI Assistants Evaluation
2 benchmarks · 1 papers
Long-term preference alignment
2 benchmarks · 1 papers
Slide Deck Generation
2 benchmarks · 3 papers
LVLM Evaluation
2 benchmarks · 2 papers
Belief Prediction
2 benchmarks · 1 papers
Human-AI Agreement Assessment
2 benchmarks · 1 papers
Hallucination Reasoning
2 benchmarks · 2 papers
Islamic inheritance reasoning
2 benchmarks · 2 papers
LLM Hallucination Detection
2 benchmarks · 3 papers
grade-school math
2 benchmarks · 1 papers
user-preference matrix generation
2 benchmarks · 1 papers
Discriminative Performance
2 benchmarks · 4 papers
Expert-level Multimodal Understanding
2 benchmarks · 1 papers
Evidence-grounded diagnostic reasoning
2 benchmarks · 1 papers
E2E Generation Latency
2 benchmarks · 1 papers
Standard Operating Procedure execution
2 benchmarks · 1 papers
Instructed Code Generation
2 benchmarks · 1 papers
General AI Assistant Task Execution
2 benchmarks · 1 papers
Uncovering hidden system prompts
2 benchmarks · 2 papers
Dialogue Simulation
2 benchmarks · 3 papers
Generative Hallucination Evaluation
2 benchmarks · 5 papers
Conversational Performance
2 benchmarks · 1 papers
Long-Horizon Search Intelligence
2 benchmarks · 2 papers
LLM Chat Evaluation
2 benchmarks · 1 papers
Human Logic Alignment
2 benchmarks · 1 papers
Leaderboard Evaluation
2 benchmarks · 1 papers
Pairwise comparison evaluation
Page 59 of 154
Previous
Next