Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Empathetic Dialogue Response Generation | 1 | 1 | Feb 21, 2026 | |
| Empathy-oriented Speech-to-Speech Dialogue | 1 | 1 | Feb 21, 2026 | |
| Language Tutoring | 1 | 1 | Apr 24, 2026 | |
| General Multitask Language Understanding | 1 | 1 | Apr 24, 2026 | |
| Contradictory-style Generation | 1 | 1 | Feb 21, 2026 | |
| Comparative Analysis of System Dimensions | 1 | 1 | Apr 24, 2026 | |
| Speech Reasoning and Response Generation | 1 | 1 | Feb 21, 2026 | |
| Get organized | 1 | 1 | Apr 24, 2026 | |
| General Knowledge Retention |
|---|
| 1 |
| 1 |
| Apr 24, 2026 |
| Storyline Generation | 1 | 1 | Apr 24, 2026 |
|---|
| Behavioral Similarity Analysis | 1 | 1 | Apr 24, 2026 |
|---|
| Question Answering Critique and Refinement | 1 | 1 | Feb 18, 2026 |
|---|
| LLM evaluation correctness | 1 | 1 | Apr 24, 2026 |
|---|
| LLM evaluation human preference | 1 | 1 | Apr 24, 2026 |
|---|
| Hard LLM Reasoning | 1 | 1 | Feb 18, 2026 |
|---|
| Calendar Scheduling | 1 | 1 | Feb 21, 2026 |
|---|
| LLM win-rate estimation ranking | 1 | 1 | Apr 24, 2026 |
|---|
| Criteria Alignment | 1 | 1 | Apr 24, 2026 |
|---|
| In-Context Reference | 1 | 1 | Apr 24, 2026 |
|---|
| Open-Ended Professional Tasks | 1 | 1 | Feb 18, 2026 |
|---|
| Mathematical logic sequence modeling | 1 | 1 | Apr 24, 2026 |
|---|
| Long-context Variable Tracking | 1 | 1 | Apr 24, 2026 |
|---|
| Reasoning and Acting | 1 | 1 | Feb 21, 2026 |
|---|
| LLM Attack Effectiveness Evaluation | 1 | 1 | Apr 24, 2026 |
|---|
| Robotic Manipulation Reasoning | 1 | 2 | Jul 3, 2026 |
|---|