Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Agent Core Capabilities Overall | 1 | 1 | Apr 20, 2026 | |
| General Conversation | 1 | 1 | Apr 20, 2026 | |
| Post-training Safety and Utility Alignment | 1 | 1 | Apr 20, 2026 | |
| Global Preference Rating | 1 | 1 | Apr 20, 2026 | |
| fact-tracing | 1 | 1 | Apr 20, 2026 | |
| Group-wise preference evaluation | 1 | 1 | Apr 20, 2026 | |
| Legal Knowledge Evaluation | 1 | 1 | Feb 21, 2026 | |
| Inference correction review (correction) | 1 | 1 | Apr 21, 2026 | |
| Multi-task Knowledge Evaluation |
| 1 |
| 1 |
| Apr 21, 2026 |
| Commonsense Story Generation | 1 | 1 | Feb 18, 2026 |
|---|
| Theory of Mind Status Prediction | 1 | 1 | Feb 18, 2026 |
|---|
| Privileged knowledge recall | 1 | 1 | Feb 18, 2026 |
|---|
| Long survey generation | 1 | 1 | Apr 21, 2026 |
|---|
| Tool-use agent scalability and performance | 1 | 1 | Apr 21, 2026 |
|---|
| Constitution Comparison | 1 | 1 | Feb 21, 2026 |
|---|
| AI-Agent Compatibility Evaluation | 1 | 1 | Apr 21, 2026 |
|---|
| Stability | 1 | 1 | Feb 21, 2026 |
|---|
| Theory of Mind Knowledge Prediction | 1 | 1 | Feb 18, 2026 |
|---|
| Zero-shot Commonsense Reasoning and Knowledge | 1 | 1 | Feb 21, 2026 |
|---|
| Error propagation analysis | 1 | 1 | Apr 21, 2026 |
|---|
| Cross-domain language model performance evaluation | 1 | 1 | Feb 21, 2026 |
|---|
| Zero-shot Evaluation Aggregate | 1 | 1 | Apr 21, 2026 |
|---|
| Online Inference | 1 | 3 | May 27, 2026 |
|---|
| Agent Selection | 1 | 1 | Apr 21, 2026 |
|---|
| Large Language Model Training Efficiency | 1 | 1 | Feb 21, 2026 |
|---|