Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
3,780 benchmarks · 2,524 papers
Mathematical Reasoning
677 benchmarks · 627 papers
Reasoning
462 benchmarks · 591 papers
Instruction Following
420 benchmarks · 178 papers
Hallucination Detection
316 benchmarks · 238 papers
Math Reasoning
228 benchmarks · 114 papers
Reward Modeling
227 benchmarks · 111 papers
Generative Modeling
163 benchmarks · 199 papers
General Reasoning
139 benchmarks · 53 papers
Knowledge Editing
131 benchmarks · 153 papers
Long-context Understanding
129 benchmarks · 196 papers
Hallucination Evaluation
123 benchmarks · 294 papers
Multi-task Language Understanding
120 benchmarks · 15 papers
LLM-generated text detection
115 benchmarks · 88 papers
Function Calling
107 benchmarks · 65 papers
Jailbreaking
104 benchmarks · 77 papers
Long-context Question Answering
103 benchmarks · 94 papers
Math
85 benchmarks · 64 papers
Dialogue Generation
80 benchmarks · 39 papers
Speculative Decoding
79 benchmarks · 117 papers
Long-context Language Understanding
79 benchmarks · 54 papers
Long-context retrieval
77 benchmarks · 172 papers
Multitask Language Understanding
76 benchmarks · 106 papers
Arithmetic Reasoning
75 benchmarks · 37 papers
Tool Calling
70 benchmarks · 65 papers
Long-context Reasoning
65 benchmarks · 19 papers
LLM Routing
63 benchmarks · 80 papers
Zero-shot Evaluation
61 benchmarks · 32 papers
Preference Alignment
56 benchmarks · 44 papers
Open-ended generation
56 benchmarks · 32 papers
Generation
54 benchmarks · 44 papers
Multi-hop Reasoning
53 benchmarks · 20 papers
Model Editing
Page 1 of 154
Previous
Next