Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Reply Generation + Tone Adjustment | 1 | 1 | Mar 13, 2026 | |
| Research Paper Reasoning and Comprehension | 1 | 1 | May 26, 2026 | |
| Language Model Evaluation Suite | 1 | 1 | Feb 19, 2026 | |
| Multi-turn math reasoning | 1 | 1 | May 26, 2026 | |
| Mathematical Verification | 1 | 1 | Jun 16, 2026 | |
| Adversarial Forgetting Evaluation | 1 | 1 | Jun 16, 2026 | |
| Multi-turn Math | 1 | 1 | May 26, 2026 | |
| Speech Generation and Interaction | 1 | 1 | Feb 19, 2026 | |
| In-distribution Tool Use |
|---|
| 1 |
| 1 |
| Mar 13, 2026 |
| LLM Jailbreak | 1 | 1 | May 26, 2026 |
|---|
| Safety intent shift | 1 | 1 | Jun 16, 2026 |
|---|
| Descriptive Question Answering | 1 | 1 | Mar 25, 2026 |
|---|
| Realworld Chat | 1 | 1 | Feb 18, 2026 |
|---|
| Persona distribution alignment with human references | 1 | 1 | Feb 19, 2026 |
|---|
| Safety alignment against harmful fine-tuning | 1 | 1 | May 26, 2026 |
|---|
| Fine-tuning Accuracy | 1 | 1 | May 26, 2026 |
|---|
| Zero-Shot Generalist Tool Use | 1 | 1 | Mar 13, 2026 |
|---|
| Model Merging for Safety and Utility | 1 | 1 | May 26, 2026 |
|---|
| Safety Boundary Over-Refusal | 1 | 1 | May 26, 2026 |
|---|
| Long-form generation hallucination evaluation | 1 | 1 | May 26, 2026 |
|---|
| Multi-Dimensional Reasoning Quality Evaluation | 1 | 1 | May 26, 2026 |
|---|
| Multimodal Task Orchestration and Question Answering | 1 | 1 | Mar 13, 2026 |
|---|
| Robust reasoning | 1 | 1 | May 26, 2026 |
|---|
| Long-context language model evaluation | 1 | 2 | Apr 13, 2026 |
|---|
| Large Language Model Downstream Evaluation | 1 | 1 | May 26, 2026 |
|---|