Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Continual Robot Manipulation | 2 | 2 | Jun 26, 2026 | |
| Multi-modal Response Generation | 2 | 1 | Feb 18, 2026 | |
| Counseling Response Generation | 1 | 1 | Feb 19, 2026 | |
| Next Instruction Prediction | 1 | 1 | Mar 17, 2026 | |
| Previous Instruction Prediction | 1 | 1 | Mar 17, 2026 | |
| Exploratory Counseling | 1 | 1 | Feb 19, 2026 | |
| Stepwise Medical Reasoning | 1 | 1 | Mar 17, 2026 | |
| Spoken Dialogue System (SDS) Semantic Quality Evaluation | 1 | 1 | Feb 19, 2026 | |
| Downstream Language Understanding and Reasoning |
|---|
| 1 |
| 1 |
| Mar 17, 2026 |
| In-Context Learning Aggregate Evaluation | 1 | 1 | Mar 17, 2026 |
|---|
| Averaged Performance across five downstream tasks | 1 | 1 | Feb 19, 2026 |
|---|
| Function-style In-Context Learning Probes | 1 | 1 | Mar 17, 2026 |
|---|
| In-Context Learning Aggregate Evaluation for Probes | 1 | 1 | Mar 17, 2026 |
|---|
| Preference Bias Mitigation | 1 | 1 | Feb 18, 2026 |
|---|
| Multilingual Reasoning and General Knowledge | 1 | 1 | Feb 19, 2026 |
|---|
| Deferral-advice | 1 | 1 | Mar 17, 2026 |
|---|
| Dialog Reasoning | 1 | 1 | Feb 18, 2026 |
|---|
| Technical problem-solving | 1 | 1 | Feb 19, 2026 |
|---|
| Narrative Generation Evaluation | 1 | 1 | Mar 17, 2026 |
|---|
| Long-context language tasks (MC, QA, Sum) | 1 | 1 | Feb 19, 2026 |
|---|
| Long-context evaluation (Financial) | 1 | 1 | Feb 19, 2026 |
|---|
| Conformal Routing Safety Control | 1 | 1 | Mar 17, 2026 |
|---|
| Moral foundation coefficient comparison in sacrificial judgments | 1 | 1 | Feb 19, 2026 |
|---|
| Spoken Scientific Reasoning | 1 | 1 | Mar 17, 2026 |
|---|
| Abstract and compositional reasoning | 1 | 1 | Mar 17, 2026 |
|---|