Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 61 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
2 benchmarks · 3 papers
Persona Simulation
2 benchmarks · 1 papers
General AI Assistant Task Execution
2 benchmarks · 1 papers
Standard Operating Procedure execution
2 benchmarks · 1 papers
Instructed Code Generation
2 benchmarks · 1 papers
Uncovering hidden system prompts
2 benchmarks · 5 papers
LLM Agent Evaluation
2 benchmarks · 3 papers
Generative Hallucination Evaluation
2 benchmarks · 2 papers
Agent Memory
2 benchmarks · 1 papers
Long-Horizon Search Intelligence
2 benchmarks · 1 papers
Evidence-grounded diagnostic reasoning
2 benchmarks · 2 papers
Multi-agent Negotiation
2 benchmarks · 3 papers
Synthetic in-context reasoning
2 benchmarks · 1 papers
Evaluation Reliability
2 benchmarks · 2 papers
Arithmetic Planning
2 benchmarks · 2 papers
Reasoning and Math
2 benchmarks · 2 papers
Multiple-choice commonsense reasoning
2 benchmarks · 3 papers
Complex Multi-step Reasoning
2 benchmarks · 2 papers
MultiModal Long-Context Understanding
2 benchmarks · 2 papers
Reward Scoring
2 benchmarks · 1 papers
High-level instruction execution
2 benchmarks · 1 papers
Alignment Reward Evaluation
2 benchmarks · 2 papers
Mathematical Calculation
2 benchmarks · 1 papers
Autoformalization and Proving
2 benchmarks · 2 papers
Zero-shot Reasoning and Question Answering
2 benchmarks · 1 papers
Runtime Controllability
2 benchmarks · 1 papers
Prompt Hygiene Evaluation
2 benchmarks · 2 papers
Helpfulness Assessment
2 benchmarks · 2 papers
Zero-shot performance evaluation
2 benchmarks · 1 papers
Multilingual LLM Evaluation
2 benchmarks · 2 papers
Visual Search and Reasoning
2 benchmarks · 1 papers
Prefill-stage hallucination risk detection
2 benchmarks · 1 papers
Large Language Model Debiasing
Page 61 of 154
Previous
Next