Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Automated CTRS evaluation | 1 | 1 | Apr 9, 2026 | |
| Multitask Language Modeling | 1 | 1 | Apr 9, 2026 | |
| Audio-Grounded Character Role-playing | 1 | 1 | Apr 9, 2026 | |
| Downstream Task Accuracy via Paradigm Routing | 1 | 1 | Apr 10, 2026 | |
| Agent Recommendation (Reflection) | 1 | 1 | Feb 18, 2026 | |
| Structured Query Instruction Following | 1 | 1 | Feb 19, 2026 | |
| Visual Instruction | 1 | 1 | Apr 10, 2026 | |
| Medical Professional Knowledge | 1 | 1 | Feb 18, 2026 | |
| Instruction-driven Image Editing and Generation |
|---|
| 1 |
| 1 |
| Apr 10, 2026 |
| End-to-end Conversational Question Answering | 1 | 1 | Apr 10, 2026 |
|---|
| Social Deduction (Undercover Role) | 1 | 1 | Feb 18, 2026 |
|---|
| Adversarial legal persuasion | 1 | 1 | Apr 10, 2026 |
|---|
| Long-context language understanding suite | 1 | 2 | Feb 21, 2026 |
|---|
| Social Deduction (Civilian Role) | 1 | 1 | Feb 18, 2026 |
|---|
| Instructional Video Editing | 1 | 1 | Feb 19, 2026 |
|---|
| Demonstration Selection | 1 | 1 | Feb 19, 2026 |
|---|
| Informal Logic Reasoning | 1 | 1 | Apr 10, 2026 |
|---|
| Large Language Modeling | 1 | 1 | Feb 18, 2026 |
|---|
| Complex instruction-based image editing | 1 | 1 | Feb 19, 2026 |
|---|
| Free-form description | 1 | 1 | Apr 10, 2026 |
|---|
| Faithful Reasoning | 1 | 1 | Feb 19, 2026 |
|---|
| Global sensemaking | 1 | 1 | Apr 10, 2026 |
|---|
| Parameter-Efficient Fine-Tuning | 1 | 1 | Feb 18, 2026 |
|---|
| Complex Multi-constraint Reasoning | 1 | 1 | Feb 19, 2026 |
|---|
| Long-context multi-task evaluation | 1 | 1 | Apr 10, 2026 |
|---|