Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Open-ended tasks | 1 | 1 | May 12, 2026 | |
| Reasoning with Latent Activations | 1 | 1 | May 12, 2026 | |
| Linguistically Diverse Reasoning | 1 | 1 | May 12, 2026 | |
| Reasoning Trajectory Generation | 1 | 1 | May 12, 2026 | |
| Logical Reasoning Question Answering | 1 | 2 | Apr 28, 2026 | |
| Idea Generation Assessment | 1 | 2 | Apr 1, 2026 | |
| Socratic Conversation Generation | 1 | 1 | May 12, 2026 | |
| Multi-hop QA Reasoning | 1 | 1 | May 12, 2026 | |
| Zero-shot Language Modeling and Commonsense Reasoning |
| 1 |
| 1 |
| May 12, 2026 |
| Semantic Task Routing | 1 | 1 | May 12, 2026 |
|---|
| Personality Expression | 1 | 1 | Mar 4, 2026 |
|---|
| Cross-domain LLM Evaluation | 1 | 1 | May 12, 2026 |
|---|
| Sales Dialogue Evaluation | 1 | 1 | Mar 4, 2026 |
|---|
| Query-based dialogue summarization | 1 | 1 | Feb 18, 2026 |
|---|
| Aggregate Out-of-Domain Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Safety Alignment Robustness Evaluation | 1 | 1 | May 12, 2026 |
|---|
| Financial Advisory Copilot | 1 | 1 | Mar 4, 2026 |
|---|
| Agent Tool-calling | 1 | 1 | May 12, 2026 |
|---|
| Personalized LLM Alignment Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Tool-use alignment | 1 | 1 | Mar 4, 2026 |
|---|
| Situation Reasoning | 1 | 1 | May 12, 2026 |
|---|
| Profanity suppression | 1 | 1 | May 12, 2026 |
|---|
| Long Text Tasks | 1 | 1 | Feb 18, 2026 |
|---|
| Long Range Arena ListOps | 1 | 1 | May 12, 2026 |
|---|
| Memory retention task | 1 | 1 | May 12, 2026 |
|---|