Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Writing Ability | 1 | 1 | Feb 18, 2026 | |
| Conversation Generation | 1 | 1 | Feb 19, 2026 | |
| Reflection Generation | 1 | 1 | Apr 3, 2026 | |
| Justification Quality Evaluation | 1 | 1 | Apr 6, 2026 | |
| Mobile-use agent task completion and intent alignment | 1 | 1 | Apr 6, 2026 | |
| Role-Play Generation | 1 | 1 | Feb 19, 2026 | |
| Expert Knowledge Q&A | 1 | 1 | Apr 6, 2026 | |
| Role-playing Reward Modeling | 1 | 1 | Feb 19, 2026 | |
| Offline Preference-Based Reinforcement Learning |
| 1 |
| 1 |
| Apr 6, 2026 |
| Unlearning Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Decision and Abstention Performance | 1 | 1 | Apr 6, 2026 |
|---|
| Prompt-based prediction ranking | 1 | 1 | Apr 6, 2026 |
|---|
| Cooperative Reasoning | 1 | 1 | Apr 6, 2026 |
|---|
| Hallucination Regeneration | 1 | 1 | Feb 18, 2026 |
|---|
| Reasoning Trajectory Evaluation | 1 | 1 | Feb 19, 2026 |
|---|
| Compositional planning | 1 | 1 | Apr 6, 2026 |
|---|
| Student vs. Teacher Win-Tie Rate evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Complex Emotion Interpretation | 1 | 1 | Feb 18, 2026 |
|---|
| Long-context retrieval/accuracy across variable lengths | 1 | 1 | Feb 19, 2026 |
|---|
| Professional deep-research writing | 1 | 1 | Apr 6, 2026 |
|---|
| Harmful Query Transformation | 1 | 1 | Feb 19, 2026 |
|---|
| URL Health and Self-Correction | 1 | 1 | Apr 6, 2026 |
|---|
| Arithmetic reasoning (multi-solution) | 1 | 1 | Feb 27, 2026 |
|---|
| Synthetic Reasoning | 1 | 1 | Feb 19, 2026 |
|---|
| Safety Resistance and Detectability Evaluation | 1 | 1 | Apr 7, 2026 |
|---|