Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Zero-shot performance evaluation | 2 | 2 | Feb 20, 2026 | |
| Multilingual LLM Evaluation | 2 | 1 | Mar 6, 2026 | |
| Belief Prediction | 2 | 2 | May 22, 2026 | |
| Arithmetic Planning | 2 | 2 | Jun 9, 2026 | |
| Complex Multi-step Reasoning | 2 | 3 | May 26, 2026 | |
| Summary Factuality Evaluation | 2 | 1 | Feb 18, 2026 | |
| Alignment Reward Evaluation | 2 | 1 | Mar 24, 2026 | |
| Social interaction simulation | 2 | 2 | Jun 9, 2026 | |
| Reward Scoring |
| 2 |
| 2 |
| Apr 6, 2026 |
| Autoformalization and Proving | 2 | 1 | Mar 23, 2026 |
|---|
| High-level instruction execution | 2 | 1 | Mar 24, 2026 |
|---|
| Prefill-stage hallucination risk detection | 2 | 1 | Mar 23, 2026 |
|---|
| Prompt Hygiene Evaluation | 2 | 1 | Mar 23, 2026 |
|---|
| Helpfulness Assessment | 2 | 2 | May 8, 2026 |
|---|
| General Language Intelligence | 2 | 2 | Apr 30, 2026 |
|---|
| Grade-school mathematical reasoning | 2 | 2 | Jun 9, 2026 |
|---|
| Large Language Model Debiasing | 2 | 1 | Mar 20, 2026 |
|---|
| Dialogue Aspect-based Sentiment Quadruple Extraction | 2 | 1 | Feb 18, 2026 |
|---|
| Generative Multiple-choice Question Answering | 2 | 1 | Feb 18, 2026 |
|---|
| Hallucination Tracing | 2 | 1 | Mar 9, 2026 |
|---|
| Audio Instruction Following | 2 | 2 | May 20, 2026 |
|---|
| Shot-Language Understanding | 2 | 1 | Mar 20, 2026 |
|---|
| Conflict Measurement | 2 | 1 | Mar 20, 2026 |
|---|
| Generative multiple-choice | 2 | 1 | Feb 18, 2026 |
|---|
| LLM Hallucination Detection | 2 | 2 | Apr 8, 2026 |
|---|