Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Dialogue Management | 4 | 2 | Feb 18, 2026 | |
| Jailbreak Transferability | 4 | 1 | Mar 4, 2026 | |
| Multimodal Knowledge Editing | 4 | 3 | May 29, 2026 | |
| Reasoning Efficiency | 3 | 2 | May 19, 2026 | |
| Linear Concept Accessibility and Steering | 3 | 1 | Apr 20, 2026 | |
| MLLM Evaluation | 3 | 2 | Mar 31, 2026 | |
| Multi-Turn Medical Dialogue | 3 | 1 | Mar 4, 2026 | |
| Reasoning trace quality evaluation | 3 | 1 | Apr 16, 2026 | |
| Tool-use Inference |
| 3 |
| 1 |
| Apr 16, 2026 |
| Process-level Reward Modeling | 3 | 1 | Mar 4, 2026 |
|---|
| Follow-up question generation | 3 | 1 | Mar 4, 2026 |
|---|
| Zero-shot Reasoning and Language Modeling | 3 | 3 | Jul 2, 2026 |
|---|
| Goal-relevance Evaluation | 3 | 1 | Apr 15, 2026 |
|---|
| Task-oriented Interaction | 3 | 1 | Mar 4, 2026 |
|---|
| Fine-grained Knowledge Recall | 3 | 1 | Apr 15, 2026 |
|---|
| Roundtrip Alignment | 3 | 1 | Feb 18, 2026 |
|---|
| Analysis Report Quality Evaluation | 3 | 1 | Mar 4, 2026 |
|---|
| MLLM-as-a-judge evaluation | 3 | 2 | Apr 21, 2026 |
|---|
| Policy Question Answering | 3 | 1 | Apr 15, 2026 |
|---|
| Multi-turn dialogue routing | 3 | 1 | Apr 15, 2026 |
|---|
| Goal selection | 3 | 1 | Mar 4, 2026 |
|---|
| Alignment defense against harmful fine-tuning | 3 | 1 | Mar 4, 2026 |
|---|
| Social Deduction Game Gameplay | 3 | 2 | May 29, 2026 |
|---|
| Long-context Input (Summarization) | 3 | 3 | Jun 4, 2026 |
|---|
| Ethical Decision-Making | 3 | 3 | May 28, 2026 |
|---|