Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 98 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
1 benchmarks · 2 papers
Free-form text generation
1 benchmarks · 1 papers
Commonsense Plausibility Estimation (Attribute Ranking)
1 benchmarks · 1 papers
Commonsense Plausibility Estimation (Free-form Text)
1 benchmarks · 1 papers
Conversational Process Model Redesign
1 benchmarks · 1 papers
Multi-discipline Multimodal Understanding and Reasoning
1 benchmarks · 1 papers
Any-to-any generation
1 benchmarks · 1 papers
cross-cultural recipe adaptation
1 benchmarks · 1 papers
Identifying preferred LLM outputs
1 benchmarks · 1 papers
Difficulty Correlation with Human Labels
1 benchmarks · 1 papers
Math and Science Reasoning
1 benchmarks · 1 papers
Difficulty Correlation with LLM Performance
1 benchmarks · 1 papers
Speech interaction S2T
1 benchmarks · 1 papers
Speech-to-text Instruction Following
1 benchmarks · 1 papers
Privacy-preserving LLM Inference
1 benchmarks · 1 papers
Difficulty Correlation with LLM Labels
1 benchmarks · 1 papers
Information Following
1 benchmarks · 1 papers
Multi-modal to Text Generation Latency
1 benchmarks · 1 papers
Domain Adaptation Utility
1 benchmarks · 1 papers
Pairwise Prediction Success Rate
1 benchmarks · 1 papers
Open-domain Reasoning
1 benchmarks · 1 papers
Metric Correlation with Steering
1 benchmarks · 1 papers
Large Language Model Serving
1 benchmarks · 1 papers
Academic Reasoning
1 benchmarks · 1 papers
Mental Health Safety Evaluation
1 benchmarks · 1 papers
Spoken Intelligence Evaluation
1 benchmarks · 1 papers
Sarcasm Explanation in Dialogue
1 benchmarks · 1 papers
Few-shot personalization and encoder-based methods evaluation
1 benchmarks · 1 papers
Reasoning Generalization
1 benchmarks · 1 papers
VLM-as-a-Judge
1 benchmarks · 1 papers
Synthetic token manipulation
1 benchmarks · 1 papers
Interactive step-by-step task guidance
1 benchmarks · 1 papers
Post-training Performance Evaluation
Page 98 of 154
Previous
Next