Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Memory-augmented Financial Agent Reasoning | 1 | 1 | Jun 2, 2026 | |
| Arabic Function Calling | 1 | 1 | Mar 19, 2026 | |
| LLM Utility Evaluation | 1 | 1 | May 20, 2026 | |
| Overton Alignment | 1 | 1 | Feb 19, 2026 | |
| Stepwise Confidence Attribution | 1 | 1 | May 20, 2026 | |
| LLM Ranking | 1 | 1 | Jun 2, 2026 | |
| Autonomous research pipeline execution | 1 | 1 | Mar 19, 2026 | |
| Aggregate Downstream Performance | 1 | 1 | Jun 2, 2026 | |
| Scattered Sequence Copying |
| 1 |
| 1 |
| Feb 18, 2026 |
| LLM Multi-Agent high-throughput serving | 1 | 1 | Jun 4, 2026 |
|---|
| Conversational Language Modeling | 1 | 1 | Mar 19, 2026 |
|---|
| Long-CoT Question Generation | 1 | 1 | Feb 19, 2026 |
|---|
| Clinical case generation | 1 | 1 | Mar 10, 2026 |
|---|
| Stepwise error detection | 1 | 1 | May 20, 2026 |
|---|
| TCM Qualitative Clinical Expert Evaluation | 1 | 1 | Mar 10, 2026 |
|---|
| Difficulty-controllable item generation | 1 | 1 | May 20, 2026 |
|---|
| Agentic Commerce | 1 | 1 | May 20, 2026 |
|---|
| Difficulty Assessment | 1 | 1 | Feb 19, 2026 |
|---|
| Decode-phase efficiency benchmarking | 1 | 1 | Mar 10, 2026 |
|---|
| Dialogue Memory | 1 | 1 | May 20, 2026 |
|---|
| Prefill Compute Efficiency | 1 | 1 | Mar 11, 2026 |
|---|
| Long-context Classification | 1 | 1 | May 20, 2026 |
|---|
| Agentic Long-context Reasoning | 1 | 1 | May 20, 2026 |
|---|
| Logical Puzzle Solving | 1 | 1 | May 20, 2026 |
|---|
| Preservation of General Capabilities | 1 | 1 | May 20, 2026 |
|---|