Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 60 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
2 benchmarks · 1 papers
General Knowledge and Instruction Following
2 benchmarks · 2 papers
Secure LLM Agent Task Completion
2 benchmarks · 1 papers
Long-term preference alignment
2 benchmarks · 8 papers
General Knowledge Assessment
2 benchmarks · 2 papers
Belief Prediction
2 benchmarks · 1 papers
general character task
2 benchmarks · 2 papers
LLM Hallucination Detection
2 benchmarks · 1 papers
Human-AI Agreement Assessment
2 benchmarks · 3 papers
grade-school math
2 benchmarks · 2 papers
Islamic inheritance reasoning
2 benchmarks · 2 papers
General Language Capabilities
2 benchmarks · 1 papers
Evidence-grounded diagnostic reasoning
2 benchmarks · 1 papers
Copy
2 benchmarks · 1 papers
General AI Assistant Task Execution
2 benchmarks · 1 papers
Standard Operating Procedure execution
2 benchmarks · 2 papers
Reasoning explanation generation
2 benchmarks · 1 papers
Long-Horizon Search Intelligence
2 benchmarks · 1 papers
Uncovering hidden system prompts
2 benchmarks · 3 papers
Generative Hallucination Evaluation
2 benchmarks · 1 papers
Instructed Code Generation
2 benchmarks · 2 papers
Multi-agent Negotiation
2 benchmarks · 1 papers
SAE latent interpretation
2 benchmarks · 2 papers
Multiple-choice commonsense reasoning
2 benchmarks · 1 papers
Evaluation Reliability
2 benchmarks · 2 papers
Mathematical Calculation
2 benchmarks · 2 papers
LLM Inference Latency
2 benchmarks · 2 papers
Zero-shot Language Modeling and Reasoning
2 benchmarks · 3 papers
Complex Multi-step Reasoning
2 benchmarks · 2 papers
MultiModal Long-Context Understanding
2 benchmarks · 2 papers
Response Moderation
2 benchmarks · 2 papers
Visual Search and Reasoning
2 benchmarks · 2 papers
Arithmetic Planning
Page 60 of 154
Previous
Next