Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| User Preference | 2 | 2 | Jul 2, 2026 | |
| Sequential Instruction Understanding | 2 | 1 | Feb 18, 2026 | |
| Rule Following | 2 | 1 | Apr 13, 2026 | |
| Long-context Multi-modal Understanding | 2 | 2 | May 14, 2026 | |
| Consultation Capability Evaluation | 2 | 1 | Apr 13, 2026 | |
| Reasoning Efficiency (Token Usage) | 2 | 1 | Feb 19, 2026 | |
| Language-conditioned simulation | 2 | 1 | Apr 14, 2026 | |
| Language Model Personalization | 2 | 1 | Feb 19, 2026 | |
| Dialog Workflow Extraction and Evaluation |
| 2 |
| 1 |
| Feb 18, 2026 |
| Long-form Question Answering with Citations | 2 | 3 | Feb 18, 2026 |
|---|
| Constraint Recall | 2 | 2 | Jul 8, 2026 |
|---|
| Language Conditioned Transfer | 2 | 1 | Apr 10, 2026 |
|---|
| Multimodal Dialogue Response Generation | 2 | 1 | Feb 19, 2026 |
|---|
| Dialogue Experience Evaluation | 2 | 1 | Feb 19, 2026 |
|---|
| LLM Serving Efficiency | 2 | 1 | Apr 10, 2026 |
|---|
| Massively Multitask Language Understanding | 2 | 2 | Mar 24, 2026 |
|---|
| Single-agent contract design | 2 | 1 | Apr 10, 2026 |
|---|
| Steering Validation | 2 | 1 | Apr 10, 2026 |
|---|
| Multi-agent contract design | 2 | 1 | Apr 10, 2026 |
|---|
| Memory Updating | 2 | 3 | Jun 25, 2026 |
|---|
| Open-ended | 2 | 5 | Apr 16, 2026 |
|---|
| Enterprise interface interaction | 2 | 1 | Apr 10, 2026 |
|---|
| Enterprise interface task completion | 2 | 1 | Apr 10, 2026 |
|---|
| Dialogue-based Multiple-choice Question Answering | 2 | 4 | Feb 18, 2026 |
|---|
| Debating | 2 | 1 | Feb 18, 2026 |
|---|