Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Large Language Model Pruning Calibration | 1 | 1 | Jun 4, 2026 | |
| Multi-round co-reference resolution | 1 | 1 | Mar 19, 2026 | |
| Long-CoT Question Generation | 1 | 1 | Feb 19, 2026 | |
| Clinical case generation | 1 | 1 | Mar 10, 2026 | |
| Stepwise error detection | 1 | 1 | May 20, 2026 | |
| Preference Adaptation | 1 | 1 | Feb 19, 2026 | |
| Chat Benchmark | 1 | 2 | Feb 18, 2026 | |
| TCM Qualitative Clinical Expert Evaluation | 1 | 1 | Mar 10, 2026 | |
| Difficulty-controllable item generation |
| 1 |
| 1 |
| May 20, 2026 |
| Generalization across multiple tasks | 1 | 1 | Jun 4, 2026 |
|---|
| Cultural Pluralism Alignment (Cultural Knowledge) | 1 | 1 | Feb 19, 2026 |
|---|
| Agentic Commerce | 1 | 1 | May 20, 2026 |
|---|
| Difficulty Assessment | 1 | 1 | Feb 19, 2026 |
|---|
| Decode-phase efficiency benchmarking | 1 | 1 | Mar 10, 2026 |
|---|
| Dialogue Memory | 1 | 1 | May 20, 2026 |
|---|
| Prefill Compute Efficiency | 1 | 1 | Mar 11, 2026 |
|---|
| Long-context Classification | 1 | 1 | May 20, 2026 |
|---|
| Agentic Long-context Reasoning | 1 | 1 | May 20, 2026 |
|---|
| Logical Puzzle Solving | 1 | 1 | May 20, 2026 |
|---|
| Preservation of General Capabilities | 1 | 1 | May 20, 2026 |
|---|
| Multi-turn Strategic Gameplay | 1 | 1 | Mar 11, 2026 |
|---|
| General Knowledge Preservation | 1 | 1 | May 20, 2026 |
|---|
| Math Word Problem Solution Generation | 1 | 2 | Feb 18, 2026 |
|---|
| Persona-based Role-Playing Faithfulness Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Conversational Intelligence | 1 | 1 | Mar 11, 2026 |
|---|