Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Complex reasoning and knowledge-based question-answering | 1 | 1 | Feb 19, 2026 | |
| Conversation labeling | 1 | 1 | Mar 31, 2026 | |
| Instruction-driven Navigation | 1 | 1 | May 27, 2026 | |
| Multilingual Long-context Classification | 1 | 1 | Feb 19, 2026 | |
| Interdisciplinary Research Ideation | 1 | 1 | Mar 13, 2026 | |
| Reasoning Step Judgment | 1 | 1 | May 27, 2026 | |
| Anchoring Effect Evaluation | 1 | 1 | Mar 31, 2026 | |
| Interdisciplinary Scientific Ideation | 1 | 1 | Mar 13, 2026 | |
| Multi-turn Coding |
|---|
| 1 |
| 1 |
| May 27, 2026 |
| Empathy evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Generation Quality and Coherence Evaluation | 1 | 1 | Feb 19, 2026 |
|---|
| Multi-turn Collaboration Editing | 1 | 1 | May 27, 2026 |
|---|
| Logical reasoning and game playing | 1 | 1 | Jun 30, 2026 |
|---|
| Mobile Application Operation | 1 | 1 | Mar 16, 2026 |
|---|
| Multi-turn Collaboration Reasoning | 1 | 1 | May 27, 2026 |
|---|
| LLM Risk Assessment | 1 | 1 | Feb 19, 2026 |
|---|
| Dangerous Knowledge Unlearning | 1 | 1 | May 27, 2026 |
|---|
| Human preference alignment for text-to-image generation | 1 | 1 | May 27, 2026 |
|---|
| Large Language Model Performance Evaluation | 1 | 1 | May 27, 2026 |
|---|
| Clinical Reasoning Evaluation | 1 | 1 | May 27, 2026 |
|---|
| Interactive Human Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Personalization Evaluation | 1 | 1 | Feb 19, 2026 |
|---|
| Multi-judge evaluation | 1 | 1 | Mar 16, 2026 |
|---|
| Multimodal capability profiling | 1 | 1 | May 27, 2026 |
|---|
| One-Time Edit | 1 | 1 | May 27, 2026 |
|---|