Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 67 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
1 benchmarks · 1 papers
In-Context Learning Aggregate Evaluation for Probes
1 benchmarks · 1 papers
Preference Bias Mitigation
1 benchmarks · 1 papers
Multilingual Reasoning and General Knowledge
1 benchmarks · 1 papers
Deferral-advice
1 benchmarks · 1 papers
Dialog Reasoning
1 benchmarks · 1 papers
Technical problem-solving
1 benchmarks · 1 papers
Narrative Generation Evaluation
1 benchmarks · 1 papers
Long-context language tasks (MC, QA, Sum)
1 benchmarks · 1 papers
Long-context evaluation (Financial)
1 benchmarks · 1 papers
Conformal Routing Safety Control
1 benchmarks · 1 papers
Moral foundation coefficient comparison in sacrificial judgments
1 benchmarks · 1 papers
Spoken Scientific Reasoning
1 benchmarks · 1 papers
Abstract and compositional reasoning
1 benchmarks · 1 papers
Zero-shot cross-task generalization
1 benchmarks · 1 papers
Single Product shopping assistance
1 benchmarks · 1 papers
Add-on Deals shopping assistance
1 benchmarks · 1 papers
Empathetic Dialogue Evaluation
1 benchmarks · 1 papers
Future Event Prediction
1 benchmarks · 1 papers
Needle In A Haystack (NIAH) video scenario
1 benchmarks · 1 papers
Long-context Information Retrieval
1 benchmarks · 1 papers
LLM Evaluation Efficiency
1 benchmarks · 1 papers
Reasoning Performance Aggregation
1 benchmarks · 1 papers
Zero-shot Downstream Reasoning
1 benchmarks · 5 papers
Assistant Response Alignment (Helpfulness and Harmlessness)
1 benchmarks · 1 papers
Alignment Quality
1 benchmarks · 1 papers
User trait simulation in agentic models
1 benchmarks · 1 papers
Mathematical function reasoning
1 benchmarks · 1 papers
Multitask Thai Language Evaluation
1 benchmarks · 1 papers
Health-domain instruction following
1 benchmarks · 1 papers
Long-context Temporal Reasoning
1 benchmarks · 1 papers
Long-form research report generation
1 benchmarks · 1 papers
completion task
Page 67 of 154
Previous
Next