Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 45 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
2 benchmarks · 1 papers
Skill Invocation Safety Auditing
2 benchmarks · 1 papers
Context Traceback
2 benchmarks · 1 papers
Multimodal Dialogue Response Generation
2 benchmarks · 1 papers
Dialogue Experience Evaluation
2 benchmarks · 1 papers
Language-conditioned simulation
2 benchmarks · 2 papers
Massively Multitask Language Understanding
2 benchmarks · 1 papers
Consultation Capability Evaluation
2 benchmarks · 1 papers
Rule Following
2 benchmarks · 5 papers
Open-ended
2 benchmarks · 4 papers
Dialogue-based Multiple-choice Question Answering
2 benchmarks · 1 papers
LLM-as-a-Judge Quality Evaluation
2 benchmarks · 1 papers
VLM-as-Judge Evaluation
2 benchmarks · 1 papers
Multi-turn Jailbreak Evaluation
2 benchmarks · 2 papers
Long-context Multi-modal Understanding
2 benchmarks · 1 papers
Debating
2 benchmarks · 1 papers
Human Auditing of Explanations
2 benchmarks · 1 papers
Molecular Science Instructions
2 benchmarks · 1 papers
Open-book generation under knowledge conflict
2 benchmarks · 1 papers
Multi-agent contract design
2 benchmarks · 1 papers
Single-agent contract design
2 benchmarks · 3 papers
Memory Updating
2 benchmarks · 1 papers
LLM Serving Efficiency
2 benchmarks · 1 papers
Steering Validation
2 benchmarks · 1 papers
Enterprise interface interaction
2 benchmarks · 1 papers
Dialogue commonsense evaluation
2 benchmarks · 1 papers
Relational Hallucination Evaluation
2 benchmarks · 1 papers
Enterprise interface task completion
2 benchmarks · 1 papers
Knowledge-to-action question answering
2 benchmarks · 1 papers
College-level Problems
2 benchmarks · 1 papers
Memory Recall Efficiency
2 benchmarks · 1 papers
Sales interaction performance
2 benchmarks · 2 papers
Game of 24
Page 45 of 154
Previous
Next