Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 89 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
1 benchmarks · 2 papers
Vision-dependent Multimodal Reasoning
1 benchmarks · 1 papers
Personalized Understanding
1 benchmarks · 1 papers
Social Deduction Game Agent Evaluation
1 benchmarks · 1 papers
Classification with expert deferral
1 benchmarks · 1 papers
Factual Correctness
1 benchmarks · 1 papers
Personalized Attribute-Reasoning Generation
1 benchmarks · 1 papers
Event Logical Reasoning
1 benchmarks · 1 papers
Instruction Hierarchy Reasoning
1 benchmarks · 1 papers
Computational Reasoning
1 benchmarks · 1 papers
Dynamic mathematical reasoning
1 benchmarks · 1 papers
Open-Form Math
1 benchmarks · 1 papers
Dense Prompt Alignment
1 benchmarks · 1 papers
All tasks combined
1 benchmarks · 1 papers
Instruction-following visual editing
1 benchmarks · 1 papers
Zero-shot Singing Voice Synthesis
1 benchmarks · 1 papers
Narrative Script Refinement
1 benchmarks · 1 papers
Media playback and ad control
1 benchmarks · 1 papers
Social media interaction
1 benchmarks · 1 papers
Cross-app workflow
1 benchmarks · 1 papers
Multi-Round Interaction Quality Evaluation
1 benchmarks · 1 papers
Tool usage in multi-turn dialogue
1 benchmarks · 1 papers
Model Card Generation
1 benchmarks · 1 papers
General Knowledge Utility
1 benchmarks · 1 papers
Story-consistent Image Generation
1 benchmarks · 1 papers
stop-or-guess task
1 benchmarks · 1 papers
Short-video comment generation
1 benchmarks · 1 papers
Visual Instruction Following Evaluation
1 benchmarks · 1 papers
Generative Insight Anticipation
1 benchmarks · 1 papers
Self-evolution
1 benchmarks · 1 papers
Smart Help
1 benchmarks · 2 papers
Reward model verification
1 benchmarks · 1 papers
Negotiation against a gpt-5.4-high-reasoning seller
Page 89 of 154
Previous
Next