Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| End-to-end Attention Latency | 1 | 1 | Apr 2, 2026 | |
| Model Merging Performance Aggregation | 1 | 1 | Feb 18, 2026 | |
| Agentic Instruction Following | 1 | 1 | Feb 19, 2026 | |
| Plausibility | 1 | 1 | Apr 2, 2026 | |
| Model Utility Maintenance | 1 | 1 | Feb 19, 2026 | |
| Intent-conditioned model selection | 1 | 1 | Apr 2, 2026 | |
| Interaction Alignment (User Study) | 1 | 1 | Feb 18, 2026 | |
| Multi-task Agent Execution | 1 | 1 | Apr 2, 2026 | |
| Jailbreak and Phishing Prompt Classification |
| 1 |
| 1 |
| Apr 2, 2026 |
| Jailbreak Prompt Classification | 1 | 1 | Apr 2, 2026 |
|---|
| Individual Alignment (User Study) | 1 | 1 | Feb 18, 2026 |
|---|
| Expert Red Teaming | 1 | 1 | Feb 19, 2026 |
|---|
| General-purpose Multimodal Understanding | 1 | 2 | May 26, 2026 |
|---|
| Offline Synthetic Data Generation | 1 | 1 | Apr 2, 2026 |
|---|
| Turn-level dialogue quality evaluation (Uses Knowledge) | 1 | 1 | Feb 18, 2026 |
|---|
| LLM Evaluation Agreement | 1 | 1 | Feb 18, 2026 |
|---|
| Phrase Protection | 1 | 1 | Feb 19, 2026 |
|---|
| Housing Consultation | 1 | 1 | Apr 2, 2026 |
|---|
| Spatial Navigation and Reasoning | 1 | 1 | Apr 2, 2026 |
|---|
| Proactive Agent Task Execution | 1 | 1 | Apr 2, 2026 |
|---|
| Multi-session Psychological Counseling | 1 | 1 | Apr 2, 2026 |
|---|
| Hallucination Evaluation (Discriminative) | 1 | 1 | Apr 2, 2026 |
|---|
| Hallucination Evaluation (Generative) | 1 | 2 | May 7, 2026 |
|---|
| Moral Drift Characterization | 1 | 1 | Apr 2, 2026 |
|---|
| LLM Judge Agreement | 1 | 1 | Feb 18, 2026 |
|---|