Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Multi-domain knowledge reasoning | 1 | 1 | Apr 1, 2026 | |
| Long-term Conversational Memory Evaluation | 1 | 1 | Apr 1, 2026 | |
| Long-running conversation | 1 | 1 | Apr 1, 2026 | |
| Long-context Conversational Memory | 1 | 1 | Apr 1, 2026 | |
| multi-turn clarification | 1 | 1 | Feb 19, 2026 | |
| Long-running Conversational Memory | 1 | 1 | Apr 1, 2026 | |
| Turn-level dialogue quality evaluation (Maintains Context) | 1 | 1 | Feb 18, 2026 | |
| Long-context compression and memory management | 1 | 1 | Apr 1, 2026 | |
| Few-shot tissue subtyping and mutation prediction |
| 1 |
| 1 |
| Feb 18, 2026 |
| Structured Data Extraction and Reasoning | 1 | 1 | Apr 1, 2026 |
|---|
| Predictive Validity Assessment | 1 | 1 | Feb 19, 2026 |
|---|
| Few-shot Biomarker and Subtype Prediction | 1 | 1 | Feb 18, 2026 |
|---|
| Human-blunder alignment | 1 | 1 | Apr 1, 2026 |
|---|
| Instruction Tracking | 1 | 1 | Feb 18, 2026 |
|---|
| Memory contextualization | 1 | 1 | Apr 1, 2026 |
|---|
| User Satisfaction Assessment | 1 | 1 | Apr 1, 2026 |
|---|
| AI-User Interaction Efficiency | 1 | 1 | Apr 1, 2026 |
|---|
| Diagnostic decision making | 1 | 1 | Apr 1, 2026 |
|---|
| Long-text patent generation | 1 | 1 | Feb 18, 2026 |
|---|
| Multi-turn conversational agent interaction | 1 | 1 | Apr 2, 2026 |
|---|
| Turn-level dialogue quality evaluation (Interesting) | 1 | 1 | Feb 18, 2026 |
|---|
| Safety Gate Capacity Scaling | 1 | 1 | Apr 2, 2026 |
|---|
| Procedural Conformance | 1 | 1 | Feb 19, 2026 |
|---|
| Safety Gate Evaluation | 1 | 1 | Apr 2, 2026 |
|---|
| Lipschitz Ball Verification | 1 | 1 | Apr 2, 2026 |
|---|