Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Large Language Model Routing and Orchestration | 1 | 1 | Apr 28, 2026 | |
| Contribution and Evidence Generation | 1 | 1 | Apr 28, 2026 | |
| Competition-level Mathematics and Science Reasoning | 1 | 1 | Feb 22, 2026 | |
| Science and general knowledge | 1 | 1 | Apr 28, 2026 | |
| Interactive LLM Alignment | 1 | 1 | Feb 18, 2026 | |
| Novel Knowledge Recall | 1 | 1 | Apr 28, 2026 | |
| Knowledge Combination | 1 | 1 | Apr 28, 2026 | |
| Role Fidelity | 1 | 1 | Feb 22, 2026 | |
| Customer Support BPM |
| 1 |
| 1 |
| Apr 28, 2026 |
| Mathematical Reasoning Critique and Refinement | 1 | 1 | Feb 18, 2026 |
|---|
| Customer Support Automation | 1 | 1 | Apr 28, 2026 |
|---|
| Copy Task | 1 | 1 | Feb 22, 2026 |
|---|
| GUI Agent Planning | 1 | 1 | Apr 28, 2026 |
|---|
| End-to-End Dialog Generation | 1 | 1 | Feb 18, 2026 |
|---|
| Financial Advisory Preference Ranking | 1 | 1 | Apr 28, 2026 |
|---|
| Attention layer forward pass | 1 | 1 | Apr 28, 2026 |
|---|
| Competence Estimation | 1 | 1 | Apr 28, 2026 |
|---|
| Harmlessness preference labeling accuracy | 1 | 1 | Feb 22, 2026 |
|---|
| Speculative Advice | 1 | 1 | Apr 28, 2026 |
|---|
| Harmfulness Refusal | 1 | 1 | Apr 28, 2026 |
|---|
| Story Premise Diversity Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Multi-step reasoning and tool use | 1 | 1 | Feb 18, 2026 |
|---|
| Data analysis step verification | 1 | 1 | Apr 28, 2026 |
|---|
| Vision Language Model Evaluation | 1 | 1 | Apr 28, 2026 |
|---|
| Speculative Decoding Throughput | 1 | 1 | Feb 27, 2026 |
|---|