Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Helpful and Harmless Response Generation | 1 | 1 | Apr 21, 2026 | |
| Generative Inference | 1 | 3 | Mar 20, 2026 | |
| Fact QA (Open Answer) | 1 | 1 | Feb 18, 2026 | |
| Correctness Calibration | 1 | 1 | Feb 21, 2026 | |
| LLM Safety Alignment | 1 | 1 | Apr 21, 2026 | |
| Preference Calibration | 1 | 1 | Feb 21, 2026 | |
| Prosocial Safety Assessment | 1 | 1 | Apr 21, 2026 | |
| Creative Writing (Story) | 1 | 1 | Feb 18, 2026 | |
| Symbolic Chain of Thought Reasoning |
| 1 |
| 1 |
| Feb 21, 2026 |
| Response Consistency Evaluation | 1 | 1 | Apr 21, 2026 |
|---|
| Agent-based Data Analysis | 1 | 1 | Feb 21, 2026 |
|---|
| Long-context Memory Modeling | 1 | 1 | Apr 21, 2026 |
|---|
| Knowledge-based Dialogue Generation | 1 | 1 | Feb 18, 2026 |
|---|
| End-to-End Dialogue Modeling | 1 | 1 | Feb 18, 2026 |
|---|
| Creative Writing (Poem) | 1 | 1 | Feb 18, 2026 |
|---|
| Steering LLM states | 1 | 1 | Feb 21, 2026 |
|---|
| Science and Engineering Question Answering | 1 | 1 | Apr 21, 2026 |
|---|
| Travel planning agent | 1 | 1 | Feb 21, 2026 |
|---|
| Multi-agent systems coordination | 1 | 1 | Apr 21, 2026 |
|---|
| Science reasoning (Welsh language) | 1 | 1 | Apr 21, 2026 |
|---|
| Table-based Mathematical Reasoning | 1 | 1 | Apr 21, 2026 |
|---|
| LLM agent alignment evaluation | 1 | 1 | Apr 21, 2026 |
|---|
| Comprehensive long-context evaluation | 1 | 1 | Apr 21, 2026 |
|---|
| Legal Knowledge and Reasoning Benchmark | 1 | 2 | Jun 9, 2026 |
|---|
| Dialogue State Transition | 1 | 1 | Feb 18, 2026 |
|---|