Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Response Text Generation | 1 | 1 | Feb 18, 2026 | |
| Medical LLM Evaluation | 1 | 2 | Mar 23, 2026 | |
| Multi-Agent Clinical Evaluation | 1 | 1 | Feb 21, 2026 | |
| Multiple-Choice Reading | 1 | 1 | Apr 23, 2026 | |
| Metacognitive Ability | 1 | 2 | Mar 27, 2026 | |
| Token Throughput | 1 | 1 | Apr 23, 2026 | |
| Long-context Text Summarization | 1 | 1 | Apr 23, 2026 | |
| Sequential Task Switching | 1 | 1 | Feb 21, 2026 | |
| Safety Alignment Verification |
| 1 |
| 1 |
| Apr 23, 2026 |
| Aggregate performance across 10 tasks | 1 | 1 | Feb 21, 2026 |
|---|
| Information Retrieval Reasoning | 1 | 1 | Apr 23, 2026 |
|---|
| Text Response Generation | 1 | 1 | Feb 18, 2026 |
|---|
| Agentic Reasoning and Interaction | 1 | 1 | Feb 18, 2026 |
|---|
| Commonsense Triple Validation | 1 | 1 | Apr 23, 2026 |
|---|
| Multi-step Visual Reasoning | 1 | 1 | Apr 23, 2026 |
|---|
| Detecting incorrect reasoning steps | 1 | 1 | Apr 23, 2026 |
|---|
| Unsafe Instruction Mitigation | 1 | 1 | Feb 21, 2026 |
|---|
| Step Verification Efficiency Evaluation | 1 | 1 | Apr 23, 2026 |
|---|
| Third-person auditing | 1 | 1 | Apr 23, 2026 |
|---|
| Proficiency Control | 1 | 1 | Feb 18, 2026 |
|---|
| Reasoning length evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Multi-session Alignment | 1 | 1 | Feb 21, 2026 |
|---|
| Reasoning Factuality | 1 | 1 | Apr 23, 2026 |
|---|
| LLM Reasoning Factuality | 1 | 1 | Apr 23, 2026 |
|---|
| Instruction-following robotic manipulation | 1 | 2 | Feb 21, 2026 |
|---|