Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Response Length Prediction | 3 | 2 | Feb 19, 2026 | |
| Reasoning accuracy | 3 | 2 | Apr 8, 2026 | |
| LLM Generation Efficiency | 3 | 1 | Feb 22, 2026 | |
| Instruction-based Video Editing | 3 | 3 | Jun 30, 2026 | |
| Bayesian Assessment of Sycophancy | 3 | 1 | Apr 21, 2026 | |
| Mathematical Reasoning Verification | 3 | 3 | May 1, 2026 | |
| Sycophancy Assessment | 3 | 1 | Apr 21, 2026 | |
| Context-aware Instruction Following | 3 | 1 | Feb 18, 2026 | |
| Linear Concept Accessibility and Steering |
| 3 |
| 1 |
| Apr 20, 2026 |
| Reasoning trace quality evaluation | 3 | 1 | Apr 16, 2026 |
|---|
| Tool-use Inference | 3 | 1 | Apr 16, 2026 |
|---|
| Competitive Mathematical Reasoning | 3 | 2 | Jun 9, 2026 |
|---|
| Output Equivalence | 3 | 1 | Feb 22, 2026 |
|---|
| Medical Instruction Following | 3 | 1 | Feb 22, 2026 |
|---|
| Pairwise LLM Judging | 3 | 2 | Jun 1, 2026 |
|---|
| Preference Labeling | 3 | 1 | Apr 21, 2026 |
|---|
| Qualitative User Preference Evaluation | 3 | 1 | Feb 18, 2026 |
|---|
| Psychological Counseling Dialogue Evaluation | 3 | 3 | May 27, 2026 |
|---|
| Attributed Text Generation | 3 | 1 | Feb 18, 2026 |
|---|
| Story Generation Evaluation | 3 | 2 | Apr 3, 2026 |
|---|
| Fine-grained Knowledge Recall | 3 | 1 | Apr 15, 2026 |
|---|
| Meeting Planning | 3 | 3 | Jun 4, 2026 |
|---|
| Multi-turn dialogue routing | 3 | 1 | Apr 15, 2026 |
|---|
| Goal-relevance Evaluation | 3 | 1 | Apr 15, 2026 |
|---|
| Multistep Soft Reasoning | 3 | 5 | Jun 30, 2026 |
|---|