Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| LLM-as-a-Judge Robustness | 2 | 1 | Feb 19, 2026 | |
| Latent multi-hop reasoning | 2 | 1 | Feb 19, 2026 | |
| Answering | 2 | 1 | Feb 18, 2026 | |
| Graduate-level Q&A | 2 | 2 | Apr 30, 2026 | |
| Complex retrieval and positional sorting | 2 | 1 | Feb 19, 2026 | |
| Free-language reasoning | 2 | 1 | Apr 10, 2026 | |
| Long Story Evaluation | 2 | 1 | Feb 19, 2026 | |
| Robot Failure Analysis (MCQ) | 2 | 1 | Apr 10, 2026 | |
| Multi-turn Editing |
|---|
| 2 |
| 2 |
| May 1, 2026 |
| Tool-Calling and Answer Generation | 2 | 1 | Feb 19, 2026 |
|---|
| Sales interaction performance | 2 | 1 | Apr 10, 2026 |
|---|
| Agent Routing | 2 | 2 | Jun 15, 2026 |
|---|
| Self-doubt detection | 2 | 1 | Apr 9, 2026 |
|---|
| Construct Validity Verification | 2 | 1 | Apr 9, 2026 |
|---|
| Helpfulness alignment | 2 | 2 | May 19, 2026 |
|---|
| Citation-augmented Question Answering | 2 | 1 | Feb 27, 2026 |
|---|
| General Agent Capability | 2 | 2 | Apr 16, 2026 |
|---|
| Harmful Question Forgetting | 2 | 1 | Apr 8, 2026 |
|---|
| Fluency Evaluation | 2 | 2 | Mar 4, 2026 |
|---|
| Pun Explanation | 2 | 1 | Apr 8, 2026 |
|---|
| Evaluator Accuracy | 2 | 2 | Apr 7, 2026 |
|---|
| Personality Recovery | 2 | 1 | Apr 8, 2026 |
|---|
| Human-Human Interaction | 2 | 2 | Apr 2, 2026 |
|---|
| Language Modeling and Question Answering | 2 | 2 | Feb 24, 2026 |
|---|
| Concept Forgetting | 2 | 1 | Apr 8, 2026 |
|---|