Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Social Interaction Evaluation | 5 | 2 | Jun 5, 2026 | |
| LLM Generation | 5 | 4 | Feb 22, 2026 | |
| Language Reasoning | 5 | 7 | Jun 19, 2026 | |
| Hallucination Robustness | 5 | 6 | Jun 19, 2026 | |
| Over-refusal | 5 | 17 | Jun 25, 2026 | |
| Reasoning failure prediction and recovery | 5 | 1 | Apr 21, 2026 | |
| Memory Question Answering | 5 | 3 | Jun 18, 2026 | |
| Deliberative Reason Index (DRI) calculation | 5 | 1 | Apr 21, 2026 | |
| Reward alignment |
| 5 |
| 3 |
| Jul 7, 2026 |
| Human-Metric Correlation | 5 | 1 | Feb 19, 2026 |
|---|
| Math and Text Reasoning | 5 | 1 | Feb 18, 2026 |
|---|
| Comprehension Questions | 5 | 1 | Apr 21, 2026 |
|---|
| Long-Context Inference | 5 | 2 | Jun 4, 2026 |
|---|
| Multi-Turn Tool Calling | 5 | 7 | Jun 25, 2026 |
|---|
| Multi-turn Dialogue Reasoning | 5 | 5 | Jun 4, 2026 |
|---|
| Qualitative Evaluation of Stance Distribution and Argument Organization | 5 | 1 | Apr 21, 2026 |
|---|
| Tool Invocation Refusal | 5 | 1 | Apr 14, 2026 |
|---|
| Sequential Composition Generalization | 5 | 1 | Apr 14, 2026 |
|---|
| Sycophancy | 5 | 4 | Jul 7, 2026 |
|---|
| Long speech understanding and reasoning | 5 | 1 | Feb 27, 2026 |
|---|
| LLM Filtering | 5 | 1 | Feb 24, 2026 |
|---|
| Process-level Evaluation | 5 | 1 | Apr 14, 2026 |
|---|
| Experience Reuse | 5 | 1 | Feb 22, 2026 |
|---|
| Zero-shot language evaluation | 5 | 5 | Jun 17, 2026 |
|---|
| Task Vector Performance | 5 | 1 | Apr 13, 2026 |
|---|