Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 78 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
1 benchmarks · 1 papers
Skill-assisted task execution
1 benchmarks · 1 papers
math Q&A
1 benchmarks · 1 papers
general multi-choice
1 benchmarks · 1 papers
Diverse reasoning tasks
1 benchmarks · 1 papers
Storytelling
1 benchmarks · 1 papers
Long Context Writing
1 benchmarks · 1 papers
Logical Consistency
1 benchmarks · 1 papers
End-to-end RAG generation
1 benchmarks · 1 papers
Generative sense-making QA
1 benchmarks · 1 papers
Textual Mathematical Reasoning
1 benchmarks · 1 papers
Multimodal Visual Logic Reasoning
1 benchmarks · 1 papers
Complex reasoning and knowledge-based question-answering
1 benchmarks · 3 papers
Human-level Standardized Exam Evaluation
1 benchmarks · 1 papers
Conversation labeling
1 benchmarks · 1 papers
Turn-level correlation with human Overall Quality ratings
1 benchmarks · 1 papers
Prompt Expansion
1 benchmarks · 1 papers
Many-shot ICL Efficiency
1 benchmarks · 1 papers
self-affirmation
1 benchmarks · 1 papers
Numerical Anchoring Effect Evaluation
1 benchmarks · 1 papers
Anchoring Effect Evaluation
1 benchmarks · 1 papers
Commonsense Reasoning and Knowledge
1 benchmarks · 1 papers
Story Customization
1 benchmarks · 1 papers
Interactive Document Generation
1 benchmarks · 1 papers
EM Decision-Making (AJSD)
1 benchmarks · 3 papers
AI Agent Reasoning and Tool-use
1 benchmarks · 1 papers
Turn-level dialogue quality evaluation (Understandable)
1 benchmarks · 1 papers
Language Modeling and Zero-shot Reasoning
1 benchmarks · 1 papers
Clinical Diagnosis and Reasoning
1 benchmarks · 2 papers
Character Consistency
1 benchmarks · 1 papers
Repetition
1 benchmarks · 1 papers
Understanding-Generation Consistency
1 benchmarks · 1 papers
Multimodal Consistency
Page 78 of 154
Previous
Next