Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Personalization Utility | 1 | 1 | Feb 18, 2026 | |
| Forward Continual Learning | 1 | 1 | May 11, 2026 | |
| Text Generation Throughput | 1 | 1 | May 11, 2026 | |
| Single-turn data-to-logic tasks | 1 | 1 | May 11, 2026 | |
| Instruction following verification | 1 | 1 | Feb 18, 2026 | |
| Harmful Prompt Refusal | 1 | 5 | Jul 3, 2026 | |
| Step-level error discrimination | 1 | 1 | May 11, 2026 | |
| Aggregate Multi-task Evaluation | 1 | 1 | Feb 25, 2026 | |
| Tool-Need Prediction |
| 1 |
| 1 |
| May 11, 2026 |
| Tool-Risk Prediction | 1 | 1 | May 11, 2026 |
|---|
| Tool-call Monitoring | 1 | 1 | May 11, 2026 |
|---|
| Tool-use runtime evaluation | 1 | 1 | May 11, 2026 |
|---|
| Long-context understanding and resolution | 1 | 1 | May 11, 2026 |
|---|
| Downstream skill execution | 1 | 1 | May 11, 2026 |
|---|
| Multimodal Cognition Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Health Reasoning | 1 | 1 | Feb 18, 2026 |
|---|
| Multi-turn planning | 1 | 1 | May 11, 2026 |
|---|
| Continual Knowledge Incorporation | 1 | 1 | May 11, 2026 |
|---|
| Fact | 1 | 1 | May 11, 2026 |
|---|
| Query routing and tool-calling accuracy evaluation | 1 | 1 | May 11, 2026 |
|---|
| LLM-Assisted Scoring | 1 | 1 | Feb 25, 2026 |
|---|
| Personalized LLM response generation | 1 | 1 | May 11, 2026 |
|---|
| Downstream Generation Impact (LLaVA-1.5) | 1 | 1 | May 11, 2026 |
|---|
| Reasoning Task | 1 | 1 | Feb 18, 2026 |
|---|
| Downstream Generation Impact (LLaVA-1.6) | 1 | 1 | May 11, 2026 |
|---|