Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Generation throughput | 3 | 1 | Feb 18, 2026 | |
| Linear Concept Accessibility and Steering | 3 | 1 | Apr 20, 2026 | |
| Long-context Input (Summarization) | 3 | 3 | Jun 4, 2026 | |
| Cognition and Reasoning | 3 | 2 | May 12, 2026 | |
| Long-context Generation (Reasoning) | 3 | 1 | Feb 27, 2026 | |
| Comprehensive Evaluation | 3 | 3 | May 28, 2026 | |
| Reward Hacking Mitigation | 3 | 1 | Feb 25, 2026 | |
| Repeated Negotiation | 3 | 1 | Feb 25, 2026 | |
| Knowledge modification |
| 3 |
| 1 |
| Feb 18, 2026 |
| Tool-use Inference | 3 | 1 | Apr 16, 2026 |
|---|
| Policy Question Answering | 3 | 1 | Apr 15, 2026 |
|---|
| Grade School Math Word Problems | 3 | 10 | Jun 5, 2026 |
|---|
| Goal-relevance Evaluation | 3 | 1 | Apr 15, 2026 |
|---|
| Instruction Awareness | 3 | 1 | Feb 18, 2026 |
|---|
| Multi-turn dialogue routing | 3 | 1 | Apr 15, 2026 |
|---|
| Multi-modal Hallucination Evaluation | 3 | 8 | Jun 30, 2026 |
|---|
| Long-form factuality evaluation | 3 | 1 | Apr 15, 2026 |
|---|
| Multi-session Retrieval-Augmented Generation | 3 | 1 | Feb 25, 2026 |
|---|
| Continual Language Modeling | 3 | 1 | Apr 27, 2026 |
|---|
| Multi-agent Cognitive Orchestration | 3 | 1 | Apr 21, 2026 |
|---|
| Reasoning trace quality evaluation | 3 | 1 | Apr 16, 2026 |
|---|
| Fine-grained Knowledge Recall | 3 | 1 | Apr 15, 2026 |
|---|
| LoRA Adapter Transfer | 3 | 1 | Apr 15, 2026 |
|---|
| Judge Agreement Accuracy | 3 | 1 | Feb 24, 2026 |
|---|
| Multi-modal Large Language Model Evaluation | 3 | 2 | May 13, 2026 |
|---|