Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Long-context downstream tasks | 1 | 1 | Feb 19, 2026 | |
| Graduate-Level Google-Proof Q&A | 1 | 1 | Mar 18, 2026 | |
| The Adding Problem | 1 | 1 | Feb 22, 2026 | |
| Holistic Social Debiasing Assessment | 1 | 1 | Feb 18, 2026 | |
| Trivia Question Answering | 1 | 1 | Mar 18, 2026 | |
| Intent Alignment and Over-personalization Detection | 1 | 1 | Feb 19, 2026 | |
| Personalized Story Generation Evaluation | 1 | 1 | Mar 18, 2026 | |
| Generative task evaluation | 1 | 1 | Feb 19, 2026 | |
| Safety Judge Performance Evaluation |
|---|
| 1 |
| 1 |
| Mar 18, 2026 |
| Prompt Optimization Efficiency | 1 | 1 | Mar 18, 2026 |
|---|
| LLM Workflow Orchestration | 1 | 1 | Mar 18, 2026 |
|---|
| IT support diagnosis and resolution | 1 | 1 | Mar 18, 2026 |
|---|
| Persona Control | 1 | 1 | Feb 27, 2026 |
|---|
| Non-Arithmetic Reasoning | 1 | 1 | Mar 18, 2026 |
|---|
| General Reasoning (Korean) | 1 | 1 | Feb 19, 2026 |
|---|
| Deep Research report generation and action adherence | 1 | 1 | Mar 18, 2026 |
|---|
| General Intelligence | 1 | 1 | Mar 18, 2026 |
|---|
| General Language Modeling Performance | 1 | 1 | Mar 18, 2026 |
|---|
| Personalized Writing | 1 | 1 | Mar 18, 2026 |
|---|
| Step Verification | 1 | 2 | May 18, 2026 |
|---|
| Multimodal Assistant Evaluation | 1 | 1 | Mar 18, 2026 |
|---|
| Human-Model Agreement Evaluation | 1 | 1 | Feb 19, 2026 |
|---|
| Generative Capability Evaluation | 1 | 1 | Mar 18, 2026 |
|---|
| Medical Response Refinement | 1 | 1 | Feb 19, 2026 |
|---|
| Cultural alignment assessment | 1 | 1 | Mar 18, 2026 |
|---|