Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Long-form RAG Evaluation | 1 | 1 | Apr 21, 2026 | |
| Prompt-based Value Alignment | 1 | 1 | Feb 18, 2026 | |
| Multi-step interaction | 1 | 1 | Feb 21, 2026 | |
| Long-form Factual Generation | 1 | 1 | Apr 21, 2026 | |
| Conversation Starter Generation | 1 | 1 | Apr 21, 2026 | |
| Real-world Reasoning | 1 | 1 | Feb 21, 2026 | |
| Conversational Starter Generation | 1 | 1 | Apr 21, 2026 | |
| Creativity Context Generation | 1 | 1 | Apr 21, 2026 | |
| Large Language Model Throughput |
| 1 |
| 1 |
| Feb 18, 2026 |
| Chat Memory Reasoning | 1 | 1 | Feb 18, 2026 |
|---|
| Creativity | 1 | 1 | Feb 21, 2026 |
|---|
| Context Generation | 1 | 1 | Apr 21, 2026 |
|---|
| Long-context conversation memory reasoning | 1 | 1 | Apr 21, 2026 |
|---|
| Closed-set Reasoning | 1 | 1 | Apr 21, 2026 |
|---|
| Long-form Ambiguous Question Answering | 1 | 1 | Apr 21, 2026 |
|---|
| mathematical deduction | 1 | 1 | Feb 18, 2026 |
|---|
| General Reasoning & QA | 1 | 1 | Feb 21, 2026 |
|---|
| Single-image Reasoning | 1 | 1 | Apr 21, 2026 |
|---|
| Dialogue flow prediction | 1 | 1 | Apr 21, 2026 |
|---|
| Tool-use and Complex Reasoning | 1 | 1 | Apr 23, 2026 |
|---|
| Human Evaluation of Personality Expression | 1 | 1 | Feb 18, 2026 |
|---|
| Long-term chatting | 1 | 1 | Feb 18, 2026 |
|---|
| Downstream Utility Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Judge Alignment | 1 | 1 | Feb 21, 2026 |
|---|
| Cultural Perspective Positioning Evaluation | 1 | 1 | Apr 23, 2026 |
|---|