Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Human Preference Evaluation | 34 | 30 | Jul 7, 2026 | |
| Conversational Recommendation | 34 | 14 | Jun 24, 2026 | |
| Human Evaluation | 33 | 30 | Jul 3, 2026 | |
| STEM Reasoning | 33 | 21 | Jul 2, 2026 | |
| Preference Prediction | 33 | 16 | Jul 3, 2026 | |
| Abductive Explanation Generation | 32 | 2 | Jun 11, 2026 | |
| Prompt Optimization | 32 | 13 | Jun 24, 2026 | |
| Alignment | 32 | 24 | Jul 7, 2026 | |
| LLM Jailbreaking |
| 32 |
| 5 |
| May 21, 2026 |
| Large Language Model Evaluation | 31 | 25 | Jun 9, 2026 |
|---|
| Insight Generation | 30 | 2 | Apr 23, 2026 |
|---|
| Detection of LLM generated text | 30 | 2 | May 12, 2026 |
|---|
| Personalization | 30 | 14 | Jul 3, 2026 |
|---|
| Instruction Tuning | 30 | 18 | Jun 4, 2026 |
|---|
| Multi-turn dialogue | 30 | 31 | Jul 7, 2026 |
|---|
| LLM-as-a-Judge | 30 | 16 | Jun 23, 2026 |
|---|
| Instruction Following Evaluation | 29 | 19 | Jun 25, 2026 |
|---|
| Selective Generation | 28 | 2 | Apr 29, 2026 |
|---|
| Personalized Text Generation | 28 | 11 | Jun 18, 2026 |
|---|
| Human Preference Alignment | 27 | 20 | Jun 9, 2026 |
|---|
| Safety Alignment Evaluation | 27 | 11 | Jul 2, 2026 |
|---|
| Knowledge Retention | 27 | 19 | Jun 24, 2026 |
|---|
| Preference Evaluation | 27 | 19 | Jun 23, 2026 |
|---|
| Multi-step Reasoning | 26 | 18 | Jul 8, 2026 |
|---|
| Sycophancy Evaluation | 26 | 14 | Jun 11, 2026 |
|---|