Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Longform generation of biographies | 1 | 1 | Feb 18, 2026 | |
| Logical reasoning multi-choice QA | 1 | 1 | Feb 18, 2026 | |
| Seeker Utterance Generation | 1 | 1 | Feb 19, 2026 | |
| Mobile Use | 1 | 1 | Mar 24, 2026 | |
| Profile Adherence | 1 | 1 | Feb 19, 2026 | |
| Domain-Specific Applied Tasks | 1 | 1 | Mar 24, 2026 | |
| Response Similarity Evaluation | 1 | 1 | Feb 18, 2026 | |
| Long-context Text Generation | 1 | 1 | Feb 19, 2026 | |
| Helpful and Harmless Preference Reasoning |
| 1 |
| 1 |
| Mar 24, 2026 |
| Debate quality evaluation alignment | 1 | 1 | Feb 18, 2026 |
|---|
| Reasoning and Shortcut Detection | 1 | 1 | Mar 24, 2026 |
|---|
| Reasoning and Classification | 1 | 2 | Apr 10, 2026 |
|---|
| Dialogue Memory Accuracy | 1 | 3 | Apr 21, 2026 |
|---|
| Multi-session collaboration | 1 | 1 | Mar 24, 2026 |
|---|
| User Simulation Intrinsic Evaluation | 1 | 1 | Mar 24, 2026 |
|---|
| Dialogue Authenticity Evaluation | 1 | 1 | Feb 27, 2026 |
|---|
| Scientific Agent Task | 1 | 1 | Mar 24, 2026 |
|---|
| Expert Evaluation of Explanations and Reasoning Trajectories | 1 | 1 | Feb 18, 2026 |
|---|
| Human-level Exams | 1 | 1 | Feb 19, 2026 |
|---|
| Preference-aware Image Generation | 1 | 1 | Mar 24, 2026 |
|---|
| Chat Preference | 1 | 2 | Mar 4, 2026 |
|---|
| Efficient Fine-tuning | 1 | 1 | Mar 24, 2026 |
|---|
| Long-form Biography Generation | 1 | 1 | Feb 18, 2026 |
|---|
| Speech Chat | 1 | 1 | Feb 18, 2026 |
|---|
| Long-context understanding and generation | 1 | 1 | Mar 24, 2026 |
|---|