Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 4 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
19 benchmarks · 6 papers
Preference Learning
19 benchmarks · 21 papers
Language Model Evaluation
19 benchmarks · 3 papers
Dialogue Quality Evaluation
19 benchmarks · 6 papers
Associative Recall
18 benchmarks · 6 papers
Long-context generation
18 benchmarks · 22 papers
General Evaluation
18 benchmarks · 3 papers
Downstream Performance Prediction
18 benchmarks · 8 papers
LLM Decoding
18 benchmarks · 8 papers
Social Reasoning
18 benchmarks · 28 papers
General Knowledge Reasoning
17 benchmarks · 9 papers
Conversational Response Generation
17 benchmarks · 8 papers
In-Context Learning
17 benchmarks · 32 papers
Truthfulness Evaluation
17 benchmarks · 17 papers
Visual Instruction Following
17 benchmarks · 15 papers
LLM Evaluation
17 benchmarks · 18 papers
Math problem solving
17 benchmarks · 2 papers
LTL Instruction Following
17 benchmarks · 11 papers
Knowledge Probing
16 benchmarks · 1 papers
Watermarking Robustness (Translation Attack)
16 benchmarks · 14 papers
Large Language Model Inference
16 benchmarks · 12 papers
Reward Model Evaluation
16 benchmarks · 17 papers
Language Understanding and Reasoning
16 benchmarks · 7 papers
Value Alignment
16 benchmarks · 5 papers
Personalized Reward Modeling
16 benchmarks · 4 papers
Hallucination Prediction
16 benchmarks · 27 papers
Multi-turn conversation
15 benchmarks · 7 papers
Activation Steering
15 benchmarks · 4 papers
Arithmetic Addition
15 benchmarks · 8 papers
Refusal Evaluation
15 benchmarks · 28 papers
General Knowledge Evaluation
15 benchmarks · 14 papers
Human preference prediction
15 benchmarks · 17 papers
General Language Understanding and Reasoning
Page 4 of 154
Previous
Next