Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Complex Factual Reasoning | 1 | 1 | May 19, 2026 | |
| Instruction Dataset Quality Evaluation | 1 | 1 | Feb 18, 2026 | |
| Blind pairwise comparison | 1 | 1 | Mar 10, 2026 | |
| Logic-heavy Reasoning | 1 | 1 | May 19, 2026 | |
| LLM Instruction Tuning | 1 | 2 | May 12, 2026 | |
| Role-Playing Ability | 1 | 1 | May 19, 2026 | |
| Multi-Turn Consistency | 1 | 1 | May 19, 2026 | |
| Theory of Mind Inference | 1 | 1 | May 19, 2026 | |
| Large Language Model Capability Evaluation |
| 1 |
| 1 |
| Jun 2, 2026 |
| Analytical and personal anecdote writing | 1 | 1 | Mar 18, 2026 |
|---|
| Prompt Robustness Evaluation | 1 | 1 | May 19, 2026 |
|---|
| Spoken Task-Oriented Dialogue | 1 | 1 | Mar 18, 2026 |
|---|
| Role-Play Faithfulness | 1 | 1 | Feb 18, 2026 |
|---|
| Memory-Augmented GUI Interaction | 1 | 1 | May 19, 2026 |
|---|
| Conversational Tool-use | 1 | 2 | May 28, 2026 |
|---|
| Generative-Evaluative Agreement | 1 | 1 | May 20, 2026 |
|---|
| Human Ranking | 1 | 1 | May 20, 2026 |
|---|
| Multi-turn Conversational Question Answering | 1 | 1 | May 20, 2026 |
|---|
| Long-context language processing | 1 | 1 | May 20, 2026 |
|---|
| Decision Making Reasoning | 1 | 1 | Mar 10, 2026 |
|---|
| LLM KV Cache Management | 1 | 1 | May 20, 2026 |
|---|
| LLM Utility Evaluation | 1 | 1 | May 20, 2026 |
|---|
| Stepwise Confidence Attribution | 1 | 1 | May 20, 2026 |
|---|
| Long-CoT Question Generation | 1 | 1 | Feb 19, 2026 |
|---|
| Clinical case generation | 1 | 1 | Mar 10, 2026 |
|---|