Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Long-form generation | 25 | 13 | Jun 9, 2026 | |
| Negotiation | 25 | 9 | Jul 8, 2026 | |
| Geometric Reasoning | 25 | 22 | Jun 25, 2026 | |
| LLM Unlearning | 25 | 15 | Jun 16, 2026 | |
| Dialogue | 24 | 21 | Jul 7, 2026 | |
| Multilingual Mathematical Reasoning | 24 | 31 | Jul 2, 2026 | |
| Prompt Classification | 23 | 7 | Jul 8, 2026 | |
| Generative Question Answering | 23 | 15 | Jun 11, 2026 | |
| Cultural safety evaluation |
| 23 |
| 1 |
| Mar 10, 2026 |
| Malicious Prompt Detection | 23 | 7 | Jul 7, 2026 |
|---|
| Knowledge-intensive reasoning | 22 | 18 | Jul 2, 2026 |
|---|
| Long-context retrieval and reasoning | 22 | 19 | Jul 1, 2026 |
|---|
| Hallucination Mitigation | 22 | 16 | Jun 16, 2026 |
|---|
| General AI Assistant Tasks | 22 | 30 | Jul 1, 2026 |
|---|
| Role-playing | 22 | 11 | Jun 26, 2026 |
|---|
| LLM-generated content detection | 21 | 3 | Jul 7, 2026 |
|---|
| Over-refusal evaluation | 21 | 18 | Jul 3, 2026 |
|---|
| World Knowledge | 21 | 11 | Jun 19, 2026 |
|---|
| Truthfulness | 20 | 34 | Jul 2, 2026 |
|---|
| Insight-level Evaluation | 20 | 1 | Apr 23, 2026 |
|---|
| Fine-tuning | 20 | 9 | Jul 7, 2026 |
|---|
| Language Model Evaluation | 19 | 21 | Jul 2, 2026 |
|---|
| Associative Recall | 19 | 6 | Jun 16, 2026 |
|---|
| Preference Learning | 19 | 6 | Apr 14, 2026 |
|---|
| Dialogue Quality Evaluation | 19 | 3 | Jun 15, 2026 |
|---|