Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Perceived system understanding | 1 | 1 | Apr 10, 2026 | |
| Process Reward Model Assessment | 1 | 3 | Apr 28, 2026 | |
| Honesty | 1 | 1 | Apr 10, 2026 | |
| Large Language Model Reasoning | 1 | 1 | Feb 18, 2026 | |
| Perspective-shift reasoning | 1 | 1 | Feb 19, 2026 | |
| Knowledgeable Deep Research | 1 | 1 | Apr 10, 2026 | |
| Instruction-following navigation | 1 | 1 | Feb 19, 2026 | |
| Long-range retrieval and reasoning | 1 | 1 | Apr 10, 2026 | |
| Holistic long-context understanding |
|---|
| 1 |
| 1 |
| Apr 10, 2026 |
| Japanese Commonsense Reasoning | 1 | 1 | Feb 18, 2026 |
|---|
| Long-Horizon Multi-Task Language Control | 1 | 1 | Feb 19, 2026 |
|---|
| Instruction-following document-level tasks | 1 | 1 | Apr 10, 2026 |
|---|
| Alignment Quality Assessment | 1 | 2 | Feb 20, 2026 |
|---|
| Visual Instruction Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Zero-shot Question Answering and Commonsense Reasoning | 1 | 2 | May 19, 2026 |
|---|
| Holistic Storytelling | 1 | 1 | Feb 19, 2026 |
|---|
| Long-term Memory Retrieval and Response Generation | 1 | 1 | Apr 10, 2026 |
|---|
| Generation and retrieval | 1 | 1 | Apr 10, 2026 |
|---|
| Agentic Memory Recall | 1 | 1 | Feb 19, 2026 |
|---|
| Instruction Following with Long-term Memory | 1 | 1 | Apr 10, 2026 |
|---|
| Safety Reasoning Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Memory-augmented Event Retrieval and Response Generation | 1 | 1 | Apr 10, 2026 |
|---|
| Instruction Realization | 1 | 1 | Apr 10, 2026 |
|---|
| Cognitive Ability Evaluation | 1 | 1 | Feb 19, 2026 |
|---|
| Knowledge Integration Quality | 1 | 1 | Apr 10, 2026 |
|---|