Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Expert Knowledge Evaluation | 1 | 1 | Feb 20, 2026 | |
| Planning and Tool Use | 1 | 1 | Apr 15, 2026 | |
| Output Specificity | 1 | 1 | Feb 20, 2026 | |
| Multi-party Dialogue Question Answering | 1 | 1 | Apr 15, 2026 | |
| Single-Hop Reasoning | 1 | 1 | Apr 15, 2026 | |
| Multi-domain Question Answering | 1 | 1 | Feb 20, 2026 | |
| Overall Reasoning (Average) | 1 | 2 | Apr 23, 2026 | |
| Factual Answering | 1 | 1 | Apr 15, 2026 | |
| Multi-Token Prediction |
| 1 |
| 1 |
| Apr 15, 2026 |
| Emotion Steering | 1 | 1 | Feb 20, 2026 |
|---|
| Validity Certification | 1 | 1 | Feb 20, 2026 |
|---|
| Multi-choice Evaluation | 1 | 1 | Apr 15, 2026 |
|---|
| Relative Robustness Analysis | 1 | 1 | Apr 15, 2026 |
|---|
| Cross-size lineage similarity detection | 1 | 1 | Apr 15, 2026 |
|---|
| Cognitive Task Assessment | 1 | 1 | Apr 15, 2026 |
|---|
| Hidden Intent Inference | 1 | 1 | Apr 15, 2026 |
|---|
| Agentic Person Search (Spatial Reasoning) | 1 | 1 | Apr 15, 2026 |
|---|
| Agentic Person Search (Temporal Reasoning) | 1 | 1 | Apr 15, 2026 |
|---|
| Knowledge-intensive language tasks evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Shuffle Dyck | 1 | 1 | Feb 27, 2026 |
|---|
| Pairwise evaluator agreement with human judgment | 1 | 1 | Apr 15, 2026 |
|---|
| Simulated Patient Portrayal | 1 | 1 | Apr 15, 2026 |
|---|
| Reward-Guided Generation | 1 | 1 | Feb 20, 2026 |
|---|
| End-to-end dialogue generation | 1 | 1 | Apr 16, 2026 |
|---|
| Reasoned Question Answering | 1 | 1 | Feb 18, 2026 |
|---|