Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Medical Long-form Answering | 3 | 1 | Feb 18, 2026 | |
| Controllable multi-objective generation | 3 | 1 | Feb 18, 2026 | |
| Social Interaction | 3 | 4 | Jul 9, 2026 | |
| Reasoning trace quality evaluation | 3 | 1 | Apr 16, 2026 | |
| Follow-up Question Answering | 3 | 1 | Feb 18, 2026 | |
| Dialogue Policy Evaluation | 3 | 1 | Feb 18, 2026 | |
| Sycophancy Assessment | 3 | 1 | Apr 21, 2026 | |
| Linear Concept Accessibility and Steering | 3 | 1 | Apr 20, 2026 | |
| Large Multimodal Model Evaluation |
| 3 |
| 10 |
| Apr 24, 2026 |
| Pairwise LLM Evaluation | 3 | 1 | Feb 18, 2026 |
|---|
| Multi-hop Question Generation | 3 | 2 | Apr 14, 2026 |
|---|
| Policy Question Answering | 3 | 1 | Apr 15, 2026 |
|---|
| Long-form Question Answering refinement | 3 | 1 | Feb 18, 2026 |
|---|
| Goal-relevance Evaluation | 3 | 1 | Apr 15, 2026 |
|---|
| Long-form generation factuality and uncertainty estimation | 3 | 1 | Feb 18, 2026 |
|---|
| Reasoning accuracy | 3 | 2 | Apr 8, 2026 |
|---|
| Bayesian Assessment of Sycophancy | 3 | 1 | Apr 21, 2026 |
|---|
| Tool-use Inference | 3 | 1 | Apr 16, 2026 |
|---|
| Fine-grained Knowledge Recall | 3 | 1 | Apr 15, 2026 |
|---|
| Output Equivalence | 3 | 1 | Feb 22, 2026 |
|---|
| Long-form factuality evaluation | 3 | 1 | Apr 15, 2026 |
|---|
| LLM Generation Efficiency | 3 | 1 | Feb 22, 2026 |
|---|
| LoRA Adapter Transfer | 3 | 1 | Apr 15, 2026 |
|---|
| Social Deduction Game Gameplay | 3 | 2 | May 29, 2026 |
|---|
| Memory Evaluation | 3 | 2 | Jun 2, 2026 |
|---|