Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 46 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
2 benchmarks · 2 papers
Long-context Multi-modal Understanding
2 benchmarks · 2 papers
Game of 24
2 benchmarks · 1 papers
General Writing
2 benchmarks · 1 papers
Rule Following
2 benchmarks · 1 papers
Language-conditioned simulation
2 benchmarks · 1 papers
Consultation Capability Evaluation
2 benchmarks · 1 papers
Multi-modal preference alignment
2 benchmarks · 1 papers
Language Conditioned Transfer
2 benchmarks · 1 papers
LLM Serving Efficiency
2 benchmarks · 1 papers
Steering Validation
2 benchmarks · 3 papers
Memory Updating
2 benchmarks · 1 papers
Single-agent contract design
2 benchmarks · 1 papers
Enterprise interface interaction
2 benchmarks · 1 papers
Enterprise interface task completion
2 benchmarks · 2 papers
Critique
2 benchmarks · 1 papers
Multi-agent contract design
2 benchmarks · 1 papers
Relational Hallucination Evaluation
2 benchmarks · 1 papers
Knowledge-to-action question answering
2 benchmarks · 1 papers
Rule-of-Thumb Generation
2 benchmarks · 3 papers
Hallucination self-detection
2 benchmarks · 2 papers
Binary Question Answering
2 benchmarks · 1 papers
Sales interaction performance
2 benchmarks · 2 papers
Graduate-level Q&A
2 benchmarks · 1 papers
Conversational Behavior Reasoning
2 benchmarks · 1 papers
Free-language reasoning
2 benchmarks · 2 papers
Language Compositionality
2 benchmarks · 1 papers
Robot Failure Analysis (MCQ)
2 benchmarks · 2 papers
Agent Routing
2 benchmarks · 1 papers
Faithfulness Hallucination Detection
2 benchmarks · 2 papers
Full-duplex Interaction
2 benchmarks · 2 papers
General Agent Capability
2 benchmarks · 1 papers
Agreement with outcome human labels
Page 46 of 154
Previous
Next