Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Memory admission | 1 | 1 | Mar 6, 2026 | |
| Oral Argument Simulation | 1 | 1 | Mar 6, 2026 | |
| General Competence | 1 | 1 | May 15, 2026 | |
| Multi-step reasoning grounded in visual information | 1 | 1 | May 15, 2026 | |
| Legal Inquisitive Dialogue | 1 | 1 | May 15, 2026 | |
| Interactive Scientific Reasoning | 1 | 1 | May 15, 2026 | |
| Direct Verifier Evaluation | 1 | 1 | May 15, 2026 | |
| Knowledge-Grounded Dialogue Generation (Fluency) | 1 | 1 | Feb 18, 2026 | |
| LaMP-1 Personalization |
| 1 |
| 2 |
| Jun 2, 2026 |
| Long-context Summary | 1 | 1 | Mar 6, 2026 |
|---|
| Finetuning with implicit harmful data | 1 | 1 | May 15, 2026 |
|---|
| Routing for Question Answering | 1 | 1 | May 15, 2026 |
|---|
| GUI Action Critiquing | 1 | 1 | May 15, 2026 |
|---|
| Multi-turn response addition | 1 | 1 | Mar 6, 2026 |
|---|
| Latent Concept Detection | 1 | 1 | May 15, 2026 |
|---|
| Multi-turn response refinement | 1 | 1 | Mar 6, 2026 |
|---|
| Commonsense Multi-hop QA | 1 | 1 | May 15, 2026 |
|---|
| LaMP-3 Personalization | 1 | 2 | Jun 2, 2026 |
|---|
| Value Alignment Text Rewriting | 1 | 1 | Mar 6, 2026 |
|---|
| Personalized Generative Recall | 1 | 1 | May 15, 2026 |
|---|
| Aggregate LLM Evaluation | 1 | 1 | May 15, 2026 |
|---|
| Multi-domain Language Understanding and Reasoning | 1 | 1 | May 15, 2026 |
|---|
| Stateful Agent-User Interaction | 1 | 1 | Mar 6, 2026 |
|---|
| Web of Lies | 1 | 1 | May 15, 2026 |
|---|
| LaMP-4 Personalization | 1 | 1 | Feb 18, 2026 |
|---|