Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Knowledge boundary assessment | 1 | 1 | Apr 30, 2026 | |
| Science Q&A | 1 | 1 | Apr 30, 2026 | |
| Overall Critique and Refinement | 1 | 1 | Feb 18, 2026 | |
| Co-Creativity Assessment | 1 | 1 | Feb 22, 2026 | |
| Incident Diagnosis and Resolution | 1 | 1 | Apr 30, 2026 | |
| Chit-chat Dialogue Generation | 1 | 1 | Feb 18, 2026 | |
| Egregious unfaithfulness | 1 | 1 | Feb 18, 2026 | |
| Multi-task Success Evaluation | 1 | 1 | Feb 22, 2026 | |
| Action-conditioned reasoning |
| 1 |
| 1 |
| Apr 30, 2026 |
| Multimodal Evaluation (Cognition/Summary) | 1 | 1 | Feb 22, 2026 |
|---|
| Value Fusion | 1 | 1 | May 1, 2026 |
|---|
| Instruction Following Safety | 1 | 1 | May 1, 2026 |
|---|
| Pairwise Safety Evaluation | 1 | 1 | May 1, 2026 |
|---|
| Long-horizon memory-based reasoning | 1 | 1 | Feb 22, 2026 |
|---|
| Instruction Hierarchy | 1 | 1 | May 1, 2026 |
|---|
| Chain Generation | 1 | 1 | Feb 18, 2026 |
|---|
| Tool-using Reasoning | 1 | 1 | Feb 18, 2026 |
|---|
| Multi-task Language Understanding and Reasoning | 1 | 1 | Feb 22, 2026 |
|---|
| Discrimination between Good Faith and Problematic agents (Peer Review) | 1 | 1 | May 1, 2026 |
|---|
| Explanation Alignment | 1 | 1 | May 1, 2026 |
|---|
| Diagnostic Dialogue | 1 | 1 | Apr 23, 2026 |
|---|
| Reasoning and Generative Tasks | 1 | 1 | Feb 27, 2026 |
|---|
| Length-controlled text generation | 1 | 1 | May 1, 2026 |
|---|
| Intent Coverage and Efficiency | 1 | 1 | May 1, 2026 |
|---|
| Best Research Idea Selection | 1 | 1 | Feb 22, 2026 |
|---|