Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Frontier Discovery and Ambiguity Disambiguation | 1 | 1 | May 18, 2026 | |
| Long-context conversation reasoning | 1 | 1 | May 18, 2026 | |
| Hallucination Steering | 1 | 1 | Mar 9, 2026 | |
| Scientific research reasoning | 1 | 1 | May 18, 2026 | |
| Emotional Dialogue Generation | 1 | 1 | Mar 10, 2026 | |
| Non-toxic generation | 1 | 1 | May 19, 2026 | |
| Professional Knowledge Reasoning | 1 | 1 | May 28, 2026 | |
| Preference Bias Mitigation | 1 | 1 | Feb 18, 2026 | |
| 1 |
| 1 |
| May 19, 2026 |
| Divergent thinking | 1 | 1 | May 28, 2026 |
|---|
| Dialog Reasoning | 1 | 1 | Feb 18, 2026 |
|---|
| General Broad Information Seeking | 1 | 1 | Feb 18, 2026 |
|---|
| Solution-Focused Brief Therapy (SFBT) | 1 | 1 | Mar 10, 2026 |
|---|
| Safety Elicitation via Prompt and Query Refinement | 1 | 1 | May 19, 2026 |
|---|
| Conversational Quality Evaluation | 1 | 1 | May 19, 2026 |
|---|
| Math Proof Reward Modeling | 1 | 1 | Feb 18, 2026 |
|---|
| Turn-level correlation with human ratings | 1 | 1 | May 19, 2026 |
|---|
| Structured Commitment Generation | 1 | 1 | May 19, 2026 |
|---|
| Winner Selection | 1 | 1 | May 19, 2026 |
|---|
| Psychiatric dialogue evaluation | 1 | 1 | Mar 10, 2026 |
|---|
| Model fidelity evaluation | 1 | 1 | May 19, 2026 |
|---|
| Multi-turn 3D Editing | 1 | 1 | May 19, 2026 |
|---|
| Mathematical Proof Reward Modeling | 1 | 1 | Feb 18, 2026 |
|---|
| Personalized Dialogue Evaluation | 1 | 1 | Mar 10, 2026 |
|---|
| Knowledge-aware refusal | 1 | 1 | May 19, 2026 |
|---|