Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Task-oriented Interaction | 3 | 1 | Mar 4, 2026 | |
| Sycophancy Assessment | 3 | 1 | Apr 21, 2026 | |
| MLLM-as-a-judge evaluation | 3 | 2 | Apr 21, 2026 | |
| Analysis Report Quality Evaluation | 3 | 1 | Mar 4, 2026 | |
| Alignment defense against harmful fine-tuning | 3 | 1 | Mar 4, 2026 | |
| Reasoning trace quality evaluation | 3 | 1 | Apr 16, 2026 | |
| Tool-use Inference | 3 | 1 | Apr 16, 2026 | |
| Long Reasoning | 3 | 1 | Mar 4, 2026 | |
| Linear Concept Accessibility and Steering |
| 3 |
| 1 |
| Apr 20, 2026 |
| Downstream evaluation | 3 | 3 | Jun 4, 2026 |
|---|
| Task-solving | 3 | 4 | May 19, 2026 |
|---|
| Goal-relevance Evaluation | 3 | 1 | Apr 15, 2026 |
|---|
| Policy Question Answering | 3 | 1 | Apr 15, 2026 |
|---|
| Memory & Reasoning | 3 | 1 | Feb 18, 2026 |
|---|
| RLHF | 3 | 4 | Jun 26, 2026 |
|---|
| Multi-session Retrieval-Augmented Generation | 3 | 1 | Feb 25, 2026 |
|---|
| Multi-turn dialogue routing | 3 | 1 | Apr 15, 2026 |
|---|
| Long-context Generation (Reasoning) | 3 | 1 | Feb 27, 2026 |
|---|
| Reward Hacking Mitigation | 3 | 1 | Feb 25, 2026 |
|---|
| Long-form factuality evaluation | 3 | 1 | Apr 15, 2026 |
|---|
| LoRA Adapter Transfer | 3 | 1 | Apr 15, 2026 |
|---|
| Social Deduction Game Gameplay | 3 | 2 | May 29, 2026 |
|---|
| Long-context Input (Summarization) | 3 | 3 | Jun 4, 2026 |
|---|
| Repeated Negotiation | 3 | 1 | Feb 25, 2026 |
|---|
| Judge Agreement Accuracy | 3 | 1 | Feb 24, 2026 |
|---|