Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Speculative Decoding Throughput | 1 | 1 | Feb 27, 2026 | |
| Empathically Neutral Compromise Generation | 1 | 1 | Apr 28, 2026 | |
| Refusal Backdoor Detection | 1 | 1 | Apr 28, 2026 | |
| Conversational Emotion Cause Extraction | 1 | 3 | Feb 18, 2026 | |
| Language Model Decoding | 1 | 1 | Feb 27, 2026 | |
| Math & Chart | 1 | 1 | Apr 28, 2026 | |
| Theory of Mind for privacy control | 1 | 1 | Feb 22, 2026 | |
| End-task performance evaluation | 1 |
| 1 |
| Apr 28, 2026 |
| Understanding Evaluation | 1 | 1 | Apr 28, 2026 |
|---|
| Agent Drift Mitigation | 1 | 1 | Apr 28, 2026 |
|---|
| Meeting Assistance | 1 | 1 | Feb 18, 2026 |
|---|
| Multi-task Knowledge and Reasoning | 1 | 4 | Jun 11, 2026 |
|---|
| Concept Monitoring | 1 | 1 | Apr 28, 2026 |
|---|
| Alignment Prediction | 1 | 1 | Feb 18, 2026 |
|---|
| Opponent Exploitation | 1 | 1 | Apr 29, 2026 |
|---|
| Observability Analysis | 1 | 1 | Jun 4, 2026 |
|---|
| Voice Interaction | 1 | 1 | Apr 29, 2026 |
|---|
| Multi-user Dialogue Response Generation | 1 | 1 | Apr 29, 2026 |
|---|
| Watermark Robustness under GPT rephrasing attack | 1 | 1 | Feb 18, 2026 |
|---|
| Dialogue Safety Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Adversarial Alignment Robustness | 1 | 1 | Feb 18, 2026 |
|---|
| Multi-turn Dialogue Alignment | 1 | 1 | Apr 29, 2026 |
|---|
| Memory construction | 1 | 1 | Feb 27, 2026 |
|---|
| General Assistant Reasoning | 1 | 1 | Feb 22, 2026 |
|---|
| Legal Multiple Choice Question Answering | 1 | 1 | Apr 29, 2026 |
|---|