Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Language Modeling and Zero-shot Multiple-Choice Reasoning | 1 | 1 | May 25, 2026 | |
| LLM hierarchy attribution | 1 | 1 | May 25, 2026 | |
| Domain Knowledge Estimation | 1 | 1 | May 25, 2026 | |
| Synthetic Dialogue Generation Evaluation | 1 | 1 | Feb 18, 2026 | |
| Proactive Sale Dialogue | 1 | 1 | Jun 9, 2026 | |
| Long-context reasoning (Pairs) | 1 | 1 | Mar 23, 2026 | |
| Overall reasoning performance | 1 | 1 | Feb 19, 2026 | |
| Alignment-diversity coverage | 1 | 1 | May 25, 2026 | |
| Truthfulness Steering |
| 1 |
| 2 |
| May 29, 2026 |
| Cognitive style steering | 1 | 1 | May 25, 2026 |
|---|
| Zero-shot Reasoning and General Knowledge | 1 | 1 | May 25, 2026 |
|---|
| Multi-turn consistency evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Input Reconstruction | 1 | 1 | May 25, 2026 |
|---|
| utterance-level pairwise preference judgement | 1 | 1 | May 25, 2026 |
|---|
| Pairwise LLM-judge evaluation | 1 | 1 | Feb 19, 2026 |
|---|
| Generative Search Policy Evaluation | 1 | 1 | Mar 12, 2026 |
|---|
| multi-turn dialogue speech evaluation | 1 | 1 | May 25, 2026 |
|---|
| RLHF Alignment Evaluation | 1 | 1 | Mar 24, 2026 |
|---|
| Conflict-resolution quality evaluation | 1 | 1 | Mar 12, 2026 |
|---|
| Open-ended Health Reasoning | 1 | 1 | May 25, 2026 |
|---|
| Aggregate General Performance | 1 | 1 | May 25, 2026 |
|---|
| Identity Internalization | 1 | 1 | May 25, 2026 |
|---|
| Watermark Embedding | 1 | 1 | May 25, 2026 |
|---|
| Long-context Factuality Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Response style evaluation | 1 | 1 | Mar 12, 2026 |
|---|