Categories
Companies
Research
Sign in
Join
Loading the SOTA2 catalog…
WizWand is now SOTA2 Research
SOTA2
Research
Papers
Benchmarks
Datasets
Tasks
LLM (Chat & Instruction) AI research – Page 116 · SOTA2 Research
Back to Research
Research domain
LLM (Chat & Instruction)
Search tasks
Search
1 benchmarks · 1 papers
Multi-agent game strategic reasoning
1 benchmarks · 1 papers
Strategic reasoning in auction games
1 benchmarks · 1 papers
Alignment with human judgment on game verification
1 benchmarks · 1 papers
Medical LLM Safety Refinement
1 benchmarks · 1 papers
Aggregate Language and Logic Tasks
1 benchmarks · 1 papers
Medical Dialogue Human-Centric Evaluation
1 benchmarks · 1 papers
Reward Modeling (Logicality)
1 benchmarks · 1 papers
Cinematic Story Generation
1 benchmarks · 1 papers
Reward Modeling (Accuracy)
1 benchmarks · 1 papers
Reward Modeling (Usefulness)
1 benchmarks · 1 papers
Continuous Story Generation
1 benchmarks · 1 papers
Instruction Alignment Evaluation
1 benchmarks · 1 papers
Malicious Goal Attack (Longer Token Generation)
1 benchmarks · 1 papers
Honesty Evaluation
1 benchmarks · 1 papers
Conversational News Recommendation
1 benchmarks · 1 papers
puzzle-4x4-task4
1 benchmarks · 1 papers
Zero-shot Multiple Choice Question Answering
1 benchmarks · 1 papers
Tool Activation Probing
1 benchmarks · 1 papers
Educational Feedback Generation
1 benchmarks · 1 papers
Abductive Logical Reasoning
1 benchmarks · 1 papers
First-Last XOR
1 benchmarks · 1 papers
Crafting ace items
1 benchmarks · 1 papers
Reasoning-intensive classification
1 benchmarks · 1 papers
Expert-level Science Reasoning
1 benchmarks · 1 papers
Grokking Detection
1 benchmarks · 1 papers
Language Modeling Utility
1 benchmarks · 1 papers
Grokking transition detection
1 benchmarks · 1 papers
General Ability Evaluation
1 benchmarks · 1 papers
Behavioral Axis Detection
1 benchmarks · 2 papers
Zero-shot tasks
1 benchmarks · 1 papers
Text-based Role Play
1 benchmarks · 1 papers
Expert Routing Consistency Analysis
Page 116 of 154
Previous
Next