Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| LLM Agent Evaluation | 2 | 5 | May 27, 2026 | |
| Multi-agent Negotiation | 2 | 2 | Apr 23, 2026 | |
| Agent Memory | 2 | 2 | Jun 19, 2026 | |
| Long-term preference alignment | 2 | 1 | Mar 27, 2026 | |
| Generative Hallucination Evaluation | 2 | 3 | May 28, 2026 | |
| Visual Search and Reasoning | 2 | 2 | May 27, 2026 | |
| Synthetic in-context reasoning | 2 | 3 | May 29, 2026 | |
| Arithmetic Planning | 2 | 2 | Jun 9, 2026 | |
| Complex Multi-step Reasoning |
| 2 |
| 3 |
| May 26, 2026 |
| Reasoning and Math | 2 | 2 | May 12, 2026 |
|---|
| High-level instruction execution | 2 | 1 | Mar 24, 2026 |
|---|
| Reward Scoring | 2 | 2 | Apr 6, 2026 |
|---|
| Alignment Reward Evaluation | 2 | 1 | Mar 24, 2026 |
|---|
| Autoformalization and Proving | 2 | 1 | Mar 23, 2026 |
|---|
| Zero-shot Reasoning and Question Answering | 2 | 2 | Apr 10, 2026 |
|---|
| Evaluation Reliability | 2 | 1 | Mar 24, 2026 |
|---|
| MultiModal Long-Context Understanding | 2 | 2 | Apr 21, 2026 |
|---|
| Prefill-stage hallucination risk detection | 2 | 1 | Mar 23, 2026 |
|---|
| Prompt Hygiene Evaluation | 2 | 1 | Mar 23, 2026 |
|---|
| Runtime Controllability | 2 | 1 | Feb 18, 2026 |
|---|
| Helpfulness Assessment | 2 | 2 | May 8, 2026 |
|---|
| Zero-shot performance evaluation | 2 | 2 | Feb 20, 2026 |
|---|
| Multilingual LLM Evaluation | 2 | 1 | Mar 6, 2026 |
|---|
| Multiple-choice commonsense reasoning | 2 | 2 | Apr 14, 2026 |
|---|
| Large Language Model Debiasing | 2 | 1 | Mar 20, 2026 |
|---|