Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 48 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
2 benchmarks · 1 papers
Retrieval-Augmented Generation (RAG)
2 benchmarks · 3 papers
Hard Reasoning
2 benchmarks · 1 papers
Conversational Numerical Reasoning Question Answering
2 benchmarks · 1 papers
Reward Modeling Suitability Evaluation
2 benchmarks · 2 papers
Graduate-level Q&A
2 benchmarks · 2 papers
Agent Routing
2 benchmarks · 1 papers
Agreement with process human labels
2 benchmarks · 1 papers
Agreement with outcome human labels
2 benchmarks · 1 papers
Free-language reasoning
2 benchmarks · 1 papers
Self-doubt detection
2 benchmarks · 1 papers
EDA
2 benchmarks · 3 papers
Language Modeling Downstream Evaluation
2 benchmarks · 1 papers
Personality Recovery
2 benchmarks · 2 papers
Multi-turn Jailbreaking
2 benchmarks · 1 papers
Harmful Question Forgetting
2 benchmarks · 1 papers
Pun Explanation
2 benchmarks · 1 papers
Factuality Hallucination
2 benchmarks · 1 papers
Concept Forgetting
2 benchmarks · 1 papers
Construct Validity Verification
2 benchmarks · 1 papers
Locality
2 benchmarks · 2 papers
Factuality Hallucination Evaluation
2 benchmarks · 1 papers
Citation Coverage Evaluation
2 benchmarks · 2 papers
Video generation reasoning
2 benchmarks · 1 papers
Robot Failure Analysis (MCQ)
2 benchmarks · 2 papers
General Agent Capability
2 benchmarks · 1 papers
Scientific Verification
2 benchmarks · 1 papers
Agentic medical interaction
2 benchmarks · 1 papers
Instruction Adherence and Security Robustness
2 benchmarks · 1 papers
Reward Model Controllability
2 benchmarks · 1 papers
Generalization to Unseen Preferences
2 benchmarks · 1 papers
Extra Task Completion
2 benchmarks · 1 papers
Agent Recommendation (Summarization)
Page 48 of 154
Previous
Next