Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Contextual | 1 | 1 | Jun 2, 2026 | |
| Long-context Instruction Following | 1 | 1 | Mar 10, 2026 | |
| Intent Mismatch Detection | 1 | 1 | May 19, 2026 | |
| Expert-Level Human Knowledge Reasoning | 1 | 1 | May 19, 2026 | |
| Deep Search and Research Reasoning | 1 | 1 | May 19, 2026 | |
| Multi-step Reasoning and Factuality | 1 | 1 | May 19, 2026 | |
| Online Troubleshooting | 1 | 1 | May 19, 2026 | |
| Correlation analysis with human preferences | 1 | 1 | Feb 18, 2026 | |
| Multi-task long-context understanding |
|---|
| 1 |
| 1 |
| Feb 19, 2026 |
| Long-horizon memory reasoning and retrieval | 1 | 1 | May 19, 2026 |
|---|
| Atomic Memory Recall | 1 | 1 | Mar 10, 2026 |
|---|
| Reasoning-chain quality evaluation | 1 | 1 | May 19, 2026 |
|---|
| Recovery Evaluation | 1 | 1 | May 19, 2026 |
|---|
| Complex Factual Reasoning | 1 | 1 | May 19, 2026 |
|---|
| Instruction Dataset Quality Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Blind pairwise comparison | 1 | 1 | Mar 10, 2026 |
|---|
| Logic-heavy Reasoning | 1 | 1 | May 19, 2026 |
|---|
| LLM Instruction Tuning | 1 | 2 | May 12, 2026 |
|---|
| Role-Playing Ability | 1 | 1 | May 19, 2026 |
|---|
| Multi-Turn Consistency | 1 | 1 | May 19, 2026 |
|---|
| Theory of Mind Inference | 1 | 1 | May 19, 2026 |
|---|
| Prompt Robustness Evaluation | 1 | 1 | May 19, 2026 |
|---|
| Role-Play Faithfulness | 1 | 1 | Feb 18, 2026 |
|---|
| Memory-Augmented GUI Interaction | 1 | 1 | May 19, 2026 |
|---|
| Conversational Tool-use | 1 | 2 | May 28, 2026 |
|---|