Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Tool-use AI Assistant | 1 | 1 | Mar 11, 2026 | |
| Qualitative evaluation of conversational responses | 1 | 1 | May 22, 2026 | |
| Reasoning and Multitask Language Understanding | 1 | 1 | Feb 19, 2026 | |
| Review Feedback Generation | 1 | 1 | Mar 11, 2026 | |
| Conversational Bandit | 1 | 1 | May 22, 2026 | |
| Multilingual Diagnostic Report Generation | 1 | 1 | Jun 8, 2026 | |
| Arithmetic Problem Solving | 1 | 2 | Feb 19, 2026 | |
| Dynamic Retrieval-Augmented Generation | 1 | 1 | May 22, 2026 |
| Multi-party cognitive stimulation dialogue | 1 | 1 | Mar 12, 2026 |
|---|
| Stage-wise response generation | 1 | 1 | May 22, 2026 |
|---|
| Cross-session memory recall | 1 | 1 | May 22, 2026 |
|---|
| Safety Jailbreak Evaluation | 1 | 1 | Mar 12, 2026 |
|---|
| Story-driven Video Generation | 1 | 1 | May 22, 2026 |
|---|
| Reasoning chain attribution | 1 | 1 | Feb 18, 2026 |
|---|
| Judge Evaluation | 1 | 1 | Mar 12, 2026 |
|---|
| Pretraining | 1 | 1 | May 22, 2026 |
|---|
| Multilingual Causal Reasoning | 1 | 1 | May 22, 2026 |
|---|
| Math Reasoning (coding tools) | 1 | 1 | Feb 19, 2026 |
|---|
| Chat Dialogue Evaluation | 1 | 1 | May 22, 2026 |
|---|
| Text-based Science Simulation | 1 | 1 | May 25, 2026 |
|---|
| Language Modeling and Zero-shot Multiple-Choice Reasoning | 1 | 1 | May 25, 2026 |
|---|
| LLM hierarchy attribution | 1 | 1 | May 25, 2026 |
|---|
| Domain Knowledge Estimation | 1 | 1 | May 25, 2026 |
|---|
| Alignment-diversity coverage | 1 | 1 | May 25, 2026 |
|---|
| Truthfulness Steering | 1 | 2 | May 29, 2026 |
|---|