Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Multi-task multimodal understanding | 1 | 1 | May 1, 2026 | |
| LLM-to-LLM Persuasion | 1 | 1 | Feb 22, 2026 | |
| Consistent Response Score | 1 | 1 | May 1, 2026 | |
| Reasoning over conflicting evidence | 1 | 1 | Feb 18, 2026 | |
| Practicality Assessment | 1 | 1 | Feb 22, 2026 | |
| Sycophancy Correction Receptiveness | 1 | 1 | May 1, 2026 | |
| Multi-round conversation | 1 | 1 | Feb 18, 2026 | |
| Grade-school-level math | 1 | 1 | Feb 18, 2026 | |
| Medical Plan Generation |
| 1 |
| 1 |
| Feb 22, 2026 |
| Continuation | 1 | 1 | May 1, 2026 |
|---|
| End-to-End Memory Performance | 1 | 1 | May 1, 2026 |
|---|
| Memory reasoning | 1 | 1 | May 1, 2026 |
|---|
| Decision-level NLL prediction | 1 | 1 | May 1, 2026 |
|---|
| Nonmonotonic reasoning | 1 | 1 | May 1, 2026 |
|---|
| Reasoning (Aggregated) | 1 | 1 | May 1, 2026 |
|---|
| Zero-shot Language Modeling and Knowledge Evaluation | 1 | 1 | May 1, 2026 |
|---|
| Multi-turn attack detection | 1 | 1 | May 1, 2026 |
|---|
| Driving with language reasoning | 1 | 1 | May 1, 2026 |
|---|
| English Reasoning | 1 | 1 | May 4, 2026 |
|---|
| Tulu generation | 1 | 1 | Feb 22, 2026 |
|---|
| Judgment Consistency | 1 | 1 | Feb 18, 2026 |
|---|
| Worst-Case Estimation Error | 1 | 1 | Feb 22, 2026 |
|---|
| Sequence-to-score reasoning | 1 | 1 | May 4, 2026 |
|---|
| Math Robustness | 1 | 1 | Feb 18, 2026 |
|---|
| Personality Adaptation | 1 | 1 | Feb 22, 2026 |
|---|