Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Long-term preference alignment | 2 | 1 | Mar 27, 2026 | |
| Secure LLM Agent Task Completion | 2 | 2 | Jun 9, 2026 | |
| Instruction Tuning Data Selection Efficiency | 2 | 2 | May 12, 2026 | |
| LLM Hallucination Detection | 2 | 2 | Apr 8, 2026 | |
| Evidence-grounded diagnostic reasoning | 2 | 1 | Mar 25, 2026 | |
| Long-sequence generation | 2 | 2 | May 22, 2026 | |
| Islamic inheritance reasoning | 2 | 2 | Apr 21, 2026 | |
| Large Language Model Watermarking | 2 |
| 2 |
| May 26, 2026 |
| Prediction Reasoning | 2 | 1 | Feb 18, 2026 |
|---|
| Out-of-Domain Reasoning | 2 | 2 | May 20, 2026 |
|---|
| Standard Operating Procedure execution | 2 | 1 | Mar 25, 2026 |
|---|
| Instructed Code Generation | 2 | 1 | Mar 25, 2026 |
|---|
| End-to-End Performance | 2 | 1 | Feb 27, 2026 |
|---|
| Query-based meeting summarization | 2 | 5 | Jun 29, 2026 |
|---|
| Analytical Reasoning | 2 | 2 | Apr 20, 2026 |
|---|
| grade-school math | 2 | 3 | Jun 30, 2026 |
|---|
| General AI Assistant Task Execution | 2 | 1 | Mar 25, 2026 |
|---|
| Malicious Goal Evaluation | 2 | 1 | Feb 18, 2026 |
|---|
| Multi-turn Retrieval-Augmented Generation | 2 | 2 | May 7, 2026 |
|---|
| Generative Hallucination Evaluation | 2 | 3 | May 28, 2026 |
|---|
| Alignment Task Evaluation | 2 | 1 | Feb 18, 2026 |
|---|
| Multi-agent Negotiation | 2 | 2 | Apr 23, 2026 |
|---|
| Mathematical Calculation | 2 | 2 | Apr 10, 2026 |
|---|
| LLM-as-Judge Response Evaluation | 2 | 1 | Feb 25, 2026 |
|---|
| Visual Search and Reasoning | 2 | 2 | May 27, 2026 |
|---|