Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| second_word_letter | 1 | 1 | Feb 18, 2026 | |
| Memory Retention Analysis | 1 | 1 | Feb 18, 2026 | |
| Context Adherence | 1 | 1 | Feb 24, 2026 | |
| End-to-End Inference Performance | 1 | 1 | May 8, 2026 | |
| Refusal behavior analysis | 1 | 1 | May 8, 2026 | |
| Multimodal Reasoning Efficiency | 1 | 1 | Feb 24, 2026 | |
| Formatting Compliance | 1 | 1 | May 8, 2026 | |
| Language Model General Utility | 1 | 1 | May 8, 2026 | |
| Nuanced Semantic Request Fulfillment |
| 1 |
| 1 |
| Feb 24, 2026 |
| Agentic Workflow Performance (Iterative Refinement Loops) | 1 | 1 | May 8, 2026 |
|---|
| Agentic Workflow Performance (Static) | 1 | 1 | May 8, 2026 |
|---|
| Distractor Effectiveness | 1 | 1 | Feb 18, 2026 |
|---|
| Sycophancy Bias Detection | 1 | 1 | Feb 18, 2026 |
|---|
| Large Language Model Inference Performance | 1 | 2 | Jun 11, 2026 |
|---|
| 14-task average | 1 | 1 | Feb 18, 2026 |
|---|
| Social Intelligence Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Long-Horizon Tool Execution | 1 | 1 | Feb 24, 2026 |
|---|
| Implicit-conflict resolution | 1 | 1 | May 8, 2026 |
|---|
| Fixed Response | 1 | 1 | May 8, 2026 |
|---|
| Root Cause Reasoning | 1 | 1 | Feb 18, 2026 |
|---|
| Dialogue-level Guidance Quality Evaluation | 1 | 1 | Feb 24, 2026 |
|---|
| Sequential Task Solving | 1 | 1 | May 8, 2026 |
|---|
| Plan-Grounded Answer Generation | 1 | 1 | Feb 24, 2026 |
|---|
| Secret loyalty evaluation | 1 | 1 | May 11, 2026 |
|---|
| Personalization Utility | 1 | 1 | Feb 18, 2026 |
|---|