Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 12 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
7 benchmarks · 1 papers
IPI Sanitization
7 benchmarks · 9 papers
Expert-Level Question Answering
7 benchmarks · 2 papers
Interactive Question Answering
7 benchmarks · 16 papers
Grade School Math Reasoning
7 benchmarks · 11 papers
Long-context language modeling evaluation
7 benchmarks · 8 papers
General Language Model Evaluation
7 benchmarks · 4 papers
Examination
7 benchmarks · 3 papers
Factual Reasoning
7 benchmarks · 4 papers
LLM Inference Performance
7 benchmarks · 7 papers
Long-context modeling
7 benchmarks · 1 papers
Knowledge & Instruction Following
7 benchmarks · 11 papers
Multi-turn Conversation Evaluation
7 benchmarks · 1 papers
Self-Healing Tool Routing
7 benchmarks · 4 papers
End-to-End Inference
7 benchmarks · 4 papers
Advanced Mathematical Reasoning
7 benchmarks · 2 papers
Decode Throughput
7 benchmarks · 5 papers
General Reasoning and Coding
7 benchmarks · 1 papers
Comparison-QA
7 benchmarks · 3 papers
Jailbreak Resistance
7 benchmarks · 2 papers
User Satisfaction Evaluation
7 benchmarks · 1 papers
Multi-turn ambiguous query clarification
7 benchmarks · 2 papers
LLM Steering
7 benchmarks · 1 papers
KV cache compression
7 benchmarks · 3 papers
Multi-Agent Reasoning
7 benchmarks · 1 papers
Time-controlled instruction-following
7 benchmarks · 1 papers
Language Modeling Inference
7 benchmarks · 3 papers
Comparison
7 benchmarks · 3 papers
Prefill Throughput
7 benchmarks · 1 papers
Concept Comprehensibility Evaluation
7 benchmarks · 11 papers
Multimodal Instruction Following
7 benchmarks · 8 papers
Question Answering and Commonsense Reasoning
7 benchmarks · 3 papers
Multi-subject customization
Page 12 of 154
Previous
Next