Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| General Knowledge Task | 1 | 1 | Feb 22, 2026 | |
| Long-range Next-token prediction | 1 | 1 | May 6, 2026 | |
| Helpful Response Evaluation | 1 | 2 | Feb 19, 2026 | |
| Refusal Induction | 1 | 1 | Feb 22, 2026 | |
| Knowledge-Orthogonal Reasoning | 1 | 1 | May 6, 2026 | |
| Role Generalization | 1 | 1 | Feb 18, 2026 | |
| Failure localization | 1 | 1 | May 6, 2026 | |
| Fact recall | 1 | 1 | May 6, 2026 | |
| LLM Agent Defense Evaluation |
| 1 |
| 1 |
| May 6, 2026 |
| Group Booking with failures (Grand Rollback) | 1 | 1 | May 6, 2026 |
|---|
| Response-Use classification | 1 | 1 | Feb 22, 2026 |
|---|
| Stealth Sycophancy Detection | 1 | 1 | May 6, 2026 |
|---|
| Zero-shot Language Modeling Evaluation | 1 | 1 | Feb 22, 2026 |
|---|
| Moral Steering | 1 | 1 | May 6, 2026 |
|---|
| Fine-grained Moral Steering | 1 | 1 | May 6, 2026 |
|---|
| Multi-turn Medical Diagnosis | 1 | 1 | Feb 18, 2026 |
|---|
| Multi-stage Reasoning and Navigation | 1 | 1 | Feb 22, 2026 |
|---|
| In-context comprehension | 1 | 1 | May 6, 2026 |
|---|
| Engagement Evaluation | 1 | 1 | Feb 22, 2026 |
|---|
| General Downstream Task Evaluation | 1 | 1 | May 6, 2026 |
|---|
| Discriminative Accuracy | 1 | 1 | May 6, 2026 |
|---|
| Perceived Authenticity Evaluation | 1 | 1 | Feb 22, 2026 |
|---|
| Downstream Policy Evaluation | 1 | 1 | May 6, 2026 |
|---|
| Overall Satisfaction Evaluation | 1 | 1 | Feb 22, 2026 |
|---|
| Long-Paragraph Factuality | 1 | 1 | Feb 18, 2026 |
|---|