Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| LLM response quality prediction | 1 | 1 | Feb 18, 2026 | |
| Tool-based multi-turn dialogue | 1 | 1 | Feb 18, 2026 | |
| Multi-hop information-seeking | 1 | 1 | Apr 8, 2026 | |
| Planning-style reasoning | 1 | 1 | Feb 19, 2026 | |
| Deep research agents / Multi-step reasoning | 1 | 1 | Apr 8, 2026 | |
| Speculative decoding evaluation | 1 | 1 | Feb 19, 2026 | |
| Lifelong Alignment | 1 | 1 | Apr 8, 2026 | |
| Structured interaction | 1 | 1 | Feb 19, 2026 | |
| Story Generation Diversity Analysis |
| 1 |
| 1 |
| Apr 8, 2026 |
| Text-centric Reasoning | 1 | 1 | Apr 8, 2026 |
|---|
| Q&A Reasoning | 1 | 1 | Apr 8, 2026 |
|---|
| Structured Response Generation | 1 | 1 | Feb 19, 2026 |
|---|
| Automated Research Paper Writing | 1 | 1 | Apr 8, 2026 |
|---|
| Multi-task Scoring | 1 | 1 | Apr 8, 2026 |
|---|
| Prior Knowledge Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Jailbreak Defense Efficiency and Effectiveness | 1 | 1 | Feb 19, 2026 |
|---|
| Subjective Rubric-based Scoring | 1 | 1 | Apr 8, 2026 |
|---|
| LLM Response Hijacking | 1 | 1 | Feb 19, 2026 |
|---|
| Uncertainty Triggering | 1 | 1 | Apr 8, 2026 |
|---|
| Multiple-choice question answering (MCQA) with evidence | 1 | 1 | Apr 8, 2026 |
|---|
| Dynamic Comprehension Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Engagement | 1 | 1 | Feb 19, 2026 |
|---|
| Medical LLM Risk Triage | 1 | 1 | Apr 8, 2026 |
|---|
| Verifiable Inference | 1 | 1 | Mar 27, 2026 |
|---|
| Multi-turn enterprise IT support diagnosis | 1 | 1 | Apr 8, 2026 |
|---|