Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Prompt-only Safety Routing | 3 | 1 | Feb 18, 2026 | |
| Prompt-Response Safety Routing | 3 | 1 | Feb 18, 2026 | |
| Safety Routing | 3 | 1 | Feb 18, 2026 | |
| Linear Concept Accessibility and Steering | 3 | 1 | Apr 20, 2026 | |
| Instruction-based Video Editing | 3 | 3 | Jun 30, 2026 | |
| Reward Hacking Mitigation | 3 | 1 | Feb 25, 2026 | |
| Reasoning trace quality evaluation | 3 | 1 | Apr 16, 2026 | |
| Lab Test Recommendation |
| 3 |
| 1 |
| Feb 27, 2026 |
| Policy Question Answering | 3 | 1 | Apr 15, 2026 |
|---|
| Goal-relevance Evaluation | 3 | 1 | Apr 15, 2026 |
|---|
| Interactive Instruction Following | 3 | 4 | Jun 26, 2026 |
|---|
| Tool-use Inference | 3 | 1 | Apr 16, 2026 |
|---|
| Judge Agreement Accuracy | 3 | 1 | Feb 24, 2026 |
|---|
| Vision-Language Instruction Following | 3 | 3 | Apr 23, 2026 |
|---|
| Fine-grained Knowledge Recall | 3 | 1 | Apr 15, 2026 |
|---|
| Multi-turn dialogue routing | 3 | 1 | Apr 15, 2026 |
|---|
| Moral Alignment | 3 | 3 | Feb 19, 2026 |
|---|
| AlpacaEval 2.0 | 3 | 1 | Feb 22, 2026 |
|---|
| Logical Refinement of Natural Language Explanations | 3 | 1 | Feb 18, 2026 |
|---|
| LoRA Adapter Transfer | 3 | 1 | Apr 15, 2026 |
|---|
| Social Deduction Game Gameplay | 3 | 2 | May 29, 2026 |
|---|
| Multi-hop reasoning and question answering | 3 | 1 | Feb 22, 2026 |
|---|
| Long-form factuality evaluation | 3 | 1 | Apr 15, 2026 |
|---|
| Output-based feature description faithfulness | 3 | 1 | Feb 18, 2026 |
|---|
| Instruction Fine-tuning | 3 | 3 | May 8, 2026 |
|---|