Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Controllability Robustness Prediction | 6 | 1 | Mar 4, 2026 | |
| LLM Performance Estimation | 6 | 2 | Jun 17, 2026 | |
| Motivational Interviewing Dialogue Evaluation | 6 | 2 | May 11, 2026 | |
| Math Word Problems | 6 | 6 | Jun 19, 2026 | |
| Long-context Summarization | 6 | 4 | May 20, 2026 | |
| Long-context answering with citations | 6 | 1 | Feb 18, 2026 | |
| Refusal Detection | 6 | 4 | May 12, 2026 | |
| Empathetic Dialogue Generation | 6 | 5 | Apr 10, 2026 |
| Multi-Task Reasoning | 6 | 7 | May 12, 2026 |
|---|
| Model Alignment | 6 | 3 | Jul 1, 2026 |
|---|
| Generative Performance | 6 | 2 | May 13, 2026 |
|---|
| Repetition Mitigation | 6 | 2 | Feb 18, 2026 |
|---|
| First Incorrect Step Identification | 6 | 1 | Feb 27, 2026 |
|---|
| Speech Instruction-Following | 6 | 3 | Jun 26, 2026 |
|---|
| Zero-shot Prediction | 6 | 3 | May 26, 2026 |
|---|
| Deep Research Evaluation | 6 | 2 | May 27, 2026 |
|---|
| Agentic Reasoning and Task Execution | 6 | 1 | Jul 8, 2026 |
|---|
| Poetry Generation | 6 | 2 | Jun 5, 2026 |
|---|
| Agentic Routing | 6 | 1 | Apr 28, 2026 |
|---|
| Multi-turn MLLM Safety Evaluation | 6 | 1 | Apr 21, 2026 |
|---|
| Long-horizon agentic task | 6 | 2 | May 26, 2026 |
|---|
| Agent Interaction | 6 | 2 | May 21, 2026 |
|---|
| LLM Watermark Spoofing | 6 | 1 | Apr 14, 2026 |
|---|
| Robot Instruction Following | 6 | 2 | Jun 1, 2026 |
|---|
| Text Completion | 5 | 1 | Apr 8, 2026 |
|---|