Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Safety Jailbreak Evaluation | 1 | 1 | Mar 12, 2026 | |
| Story-driven Video Generation | 1 | 1 | May 22, 2026 | |
| Reasoning chain attribution | 1 | 1 | Feb 18, 2026 | |
| Judge Evaluation | 1 | 1 | Mar 12, 2026 | |
| Pretraining | 1 | 1 | May 22, 2026 | |
| Diverse Reasoning | 1 | 1 | Jun 9, 2026 | |
| Long-generation | 1 | 1 | Mar 23, 2026 | |
| Open-ended refusal | 1 | 1 | Jun 9, 2026 | |
| Multilingual Causal Reasoning |
| 1 |
| 1 |
| May 22, 2026 |
| Math Reasoning (coding tools) | 1 | 1 | Feb 19, 2026 |
|---|
| Chat Dialogue Evaluation | 1 | 1 | May 22, 2026 |
|---|
| Psychotherapy Dialogue Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Text-based Science Simulation | 1 | 1 | May 25, 2026 |
|---|
| Language Modeling and Zero-shot Multiple-Choice Reasoning | 1 | 1 | May 25, 2026 |
|---|
| LLM hierarchy attribution | 1 | 1 | May 25, 2026 |
|---|
| Domain Knowledge Estimation | 1 | 1 | May 25, 2026 |
|---|
| Alignment-diversity coverage | 1 | 1 | May 25, 2026 |
|---|
| Truthfulness Steering | 1 | 2 | May 29, 2026 |
|---|
| Cognitive style steering | 1 | 1 | May 25, 2026 |
|---|
| Zero-shot Reasoning and General Knowledge | 1 | 1 | May 25, 2026 |
|---|
| Multi-turn consistency evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Input Reconstruction | 1 | 1 | May 25, 2026 |
|---|
| utterance-level pairwise preference judgement | 1 | 1 | May 25, 2026 |
|---|
| Pairwise LLM-judge evaluation | 1 | 1 | Feb 19, 2026 |
|---|
| Generative Search Policy Evaluation | 1 | 1 | Mar 12, 2026 |
|---|