Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 101 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
1 benchmarks · 1 papers
Fermi Problem Solving
1 benchmarks · 1 papers
Multitask Evaluation
1 benchmarks · 1 papers
Medical AI Performance Evaluation
1 benchmarks · 1 papers
General Assistant Task Solving
1 benchmarks · 1 papers
Autonomous Agent Problem Solving
1 benchmarks · 1 papers
Patient Simulation Fidelity
1 benchmarks · 1 papers
Multi-hop Reasoning and Fact-checking
1 benchmarks · 1 papers
General Assistant Reasoning
1 benchmarks · 1 papers
Simulated decision-making
1 benchmarks · 1 papers
Multi-turn Dialogue Quality
1 benchmarks · 1 papers
Zero-shot Language Reasoning
1 benchmarks · 6 papers
Truthful QA
1 benchmarks · 1 papers
Dialog Critique and Refinement
1 benchmarks · 1 papers
Intent-Based Interaction Generation
1 benchmarks · 1 papers
multi-turn interaction-based problem solving
1 benchmarks · 1 papers
General Reasoning Generalization
1 benchmarks · 1 papers
Agent Communication Language Evaluation
1 benchmarks · 1 papers
Social Deduction Game
1 benchmarks · 1 papers
Pedagogical Dialogue Classification
1 benchmarks · 1 papers
Multi-edit Instruction Adherence
1 benchmarks · 1 papers
Autonomous LLM Agent Verification
1 benchmarks · 1 papers
Collaborative decision-making
1 benchmarks · 1 papers
Ablation hypothesis quality evaluation
1 benchmarks · 1 papers
Malicious Requestor Interaction Selection
1 benchmarks · 1 papers
Collaborative text generation
1 benchmarks · 1 papers
Interactive Mathematical Reasoning
1 benchmarks · 1 papers
Response Text Generation
1 benchmarks · 2 papers
Medical LLM Evaluation
1 benchmarks · 1 papers
Multi-Agent Clinical Evaluation
1 benchmarks · 1 papers
Multiple-Choice Reading
1 benchmarks · 2 papers
Metacognitive Ability
1 benchmarks · 1 papers
Token Throughput
Page 101 of 154
Previous
Next