Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 125 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
1 benchmarks · 1 papers
Safety certification of tool-integrated LLMs
1 benchmarks · 1 papers
Multitask Knowledge Evaluation
1 benchmarks · 1 papers
Multi-turn psychiatric consultation
1 benchmarks · 1 papers
Persuasive essay production against scientific consensus
1 benchmarks · 1 papers
Support-faithfulness evaluation
1 benchmarks · 1 papers
Tool-use Agent Tasks
1 benchmarks · 1 papers
Intent-conditioned counterspeech generation
1 benchmarks · 1 papers
Conversational Memory Question Answering
1 benchmarks · 2 papers
Automated Interpretability Evaluation
1 benchmarks · 1 papers
Commonsense Reasoning and Short-Context Language Understanding
1 benchmarks · 1 papers
Multidisciplinary knowledge and reasoning
1 benchmarks · 1 papers
Encoding Speed
1 benchmarks · 1 papers
Hallucination Suppression
1 benchmarks · 1 papers
Multi-hop Multiple-choice Question Answering
1 benchmarks · 1 papers
Refusal Ablation
1 benchmarks · 1 papers
Psychological Assessment Game Generation Evaluation
1 benchmarks · 1 papers
Steering Success Rate
1 benchmarks · 1 papers
Helpful Assistant
1 benchmarks · 1 papers
Maritime Dialogue Generation
1 benchmarks · 1 papers
Commonsense Reasoning and Question Answering
1 benchmarks · 1 papers
Prompt-to-prompt semantic composition
1 benchmarks · 1 papers
Instruction Following and Tool Use
1 benchmarks · 1 papers
Long-Term Credit Assignment
1 benchmarks · 1 papers
Ranking LLM solutions
1 benchmarks · 1 papers
Sequential Reasoning
1 benchmarks · 1 papers
Social dilemma resolution
1 benchmarks · 1 papers
mathematical reasoning via verifiable subquestions
1 benchmarks · 1 papers
Human Ranking of Explanation Quality
1 benchmarks · 1 papers
Open-ended generative Question Answering
1 benchmarks · 1 papers
Time to First Token
1 benchmarks · 1 papers
Trade Negotiation
1 benchmarks · 1 papers
research-level mathematical reasoning
Page 125 of 154
Previous
Next