Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Safety alignment against harmful fine-tuning | 1 | 1 | May 26, 2026 | |
| Fine-tuning Accuracy | 1 | 1 | May 26, 2026 | |
| Zero-Shot Generalist Tool Use | 1 | 1 | Mar 13, 2026 | |
| Model Merging for Safety and Utility | 1 | 1 | May 26, 2026 | |
| Safety Boundary Over-Refusal | 1 | 1 | May 26, 2026 | |
| Long-form generation hallucination evaluation | 1 | 1 | May 26, 2026 | |
| Physical Reasoning and Instruction Following | 1 | 1 | Jun 16, 2026 | |
| Checkpoint Selection |
| 1 |
| 1 |
| Mar 25, 2026 |
| Multi-Dimensional Reasoning Quality Evaluation | 1 | 1 | May 26, 2026 |
|---|
| Multimodal Task Orchestration and Question Answering | 1 | 1 | Mar 13, 2026 |
|---|
| Robust reasoning | 1 | 1 | May 26, 2026 |
|---|
| State Updating | 1 | 1 | Jun 17, 2026 |
|---|
| Code-related memory dialogue | 1 | 1 | Mar 25, 2026 |
|---|
| Long-context language model evaluation | 1 | 2 | Apr 13, 2026 |
|---|
| Large Language Model Downstream Evaluation | 1 | 1 | May 26, 2026 |
|---|
| Multi-turn conversational quality | 1 | 1 | May 26, 2026 |
|---|
| Writing capabilities | 1 | 1 | May 26, 2026 |
|---|
| Service and Safety Evaluation | 1 | 1 | May 26, 2026 |
|---|
| Memorized Content Extraction | 1 | 1 | Mar 13, 2026 |
|---|
| Online Customer Support | 1 | 1 | May 26, 2026 |
|---|
| LLM Workflow Optimization | 1 | 1 | Feb 18, 2026 |
|---|
| Online Livestream Interaction | 1 | 1 | May 26, 2026 |
|---|
| Disagreement Handling | 1 | 1 | Mar 13, 2026 |
|---|
| Mathematical Reasoning Repair | 1 | 1 | May 26, 2026 |
|---|
| Routine Task Management | 1 | 1 | Mar 13, 2026 |
|---|