Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Agent Safety Reasoning | 1 | 1 | Apr 7, 2026 | |
| General Large Language Model Evaluation | 1 | 2 | Feb 25, 2026 | |
| General Language Performance | 1 | 1 | Feb 19, 2026 | |
| Scaling Efficiency | 1 | 1 | Apr 7, 2026 | |
| Steered Language Generation | 1 | 1 | Apr 7, 2026 | |
| Persona-based Dialogue | 1 | 1 | Feb 19, 2026 | |
| Arithmetic Errors (AE) Vulnerability Detection | 1 | 1 | Apr 7, 2026 | |
| Dynamic Product Ads Ranking | 1 | 1 | Apr 7, 2026 | |
| General Multimodal Intelligence Evaluation |
|---|
| 1 |
| 1 |
| Feb 18, 2026 |
| Writing capability evaluation | 1 | 1 | Feb 19, 2026 |
|---|
| Hallucination Localization | 1 | 1 | Apr 7, 2026 |
|---|
| Harmful Content Refusal | 1 | 1 | Feb 19, 2026 |
|---|
| Expert Medical Knowledge MCQ | 1 | 1 | Apr 7, 2026 |
|---|
| General Preference Pairwise | 1 | 1 | Apr 7, 2026 |
|---|
| Medical Chat Evaluation | 1 | 1 | Apr 7, 2026 |
|---|
| Open-Ended Medical Evaluation | 1 | 1 | Apr 7, 2026 |
|---|
| Multimodal Multi-task Understanding | 1 | 1 | Apr 7, 2026 |
|---|
| Helpful Assistants Alignment | 1 | 1 | Apr 7, 2026 |
|---|
| Value Control | 1 | 1 | Feb 18, 2026 |
|---|
| Image-Grounded Dialogue Generation | 1 | 1 | Feb 27, 2026 |
|---|
| Dialogue Evaluation Human Correlation | 1 | 2 | Feb 18, 2026 |
|---|
| Personal Question Answering | 1 | 1 | Feb 18, 2026 |
|---|
| Head-to-Head Comparative Evaluation | 1 | 1 | Feb 19, 2026 |
|---|
| General QA / Reasoning | 1 | 1 | Apr 7, 2026 |
|---|
| Skill Specification Quality Assessment | 1 | 1 | Apr 7, 2026 |
|---|