Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Dangerous Knowledge Unlearning | 1 | 1 | May 27, 2026 | |
| Human preference alignment for text-to-image generation | 1 | 1 | May 27, 2026 | |
| Large Language Model Performance Evaluation | 1 | 1 | May 27, 2026 | |
| Clinical Reasoning Evaluation | 1 | 1 | May 27, 2026 | |
| Interactive Human Evaluation | 1 | 1 | Feb 18, 2026 | |
| Personalization Evaluation | 1 | 1 | Feb 19, 2026 | |
| Multi-judge evaluation | 1 | 1 | Mar 16, 2026 | |
| Multimodal capability profiling |
| 1 |
| 1 |
| May 27, 2026 |
| Long-term Conversational Memory Evaluation | 1 | 1 | Apr 1, 2026 |
|---|
| Logic Inversion | 1 | 1 | Jul 2, 2026 |
|---|
| Multi-turn escalation defense | 1 | 1 | Jul 3, 2026 |
|---|
| Long-context Conversational Memory | 1 | 1 | Apr 1, 2026 |
|---|
| Helpful Assistant task (Helpfulness and Harmlessness) | 1 | 1 | Jul 3, 2026 |
|---|
| Turn-level dialogue quality evaluation (Maintains Context) | 1 | 1 | Feb 18, 2026 |
|---|
| One-Time Edit | 1 | 1 | May 27, 2026 |
|---|
| Short-chain composition | 1 | 1 | May 27, 2026 |
|---|
| Long-context Memory Retrieval | 1 | 9 | Jun 12, 2026 |
|---|
| General Reasoning and Creative Writing | 1 | 1 | Mar 16, 2026 |
|---|
| Context Learning Task-Solving | 1 | 1 | May 27, 2026 |
|---|
| Personalized Memory Retrieval | 1 | 1 | Feb 19, 2026 |
|---|
| LLM Routing Resource Efficiency | 1 | 1 | Mar 16, 2026 |
|---|
| Human Evaluation of Value Alignment | 1 | 1 | May 27, 2026 |
|---|
| Therapeutic Competence Evaluation | 1 | 1 | Feb 19, 2026 |
|---|
| Proactive Task Management | 1 | 1 | May 27, 2026 |
|---|
| Interleaved Reactive-Proactive Requests | 1 | 1 | May 27, 2026 |
|---|