Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Large Language Model Serving | 1 | 1 | Apr 21, 2026 | |
| Academic Reasoning | 1 | 1 | Feb 21, 2026 | |
| Mental Health Safety Evaluation | 1 | 1 | Apr 21, 2026 | |
| Spoken Intelligence Evaluation | 1 | 1 | Feb 18, 2026 | |
| Sarcasm Explanation in Dialogue | 1 | 1 | Feb 18, 2026 | |
| Few-shot personalization and encoder-based methods evaluation | 1 | 1 | Feb 18, 2026 | |
| Reasoning Generalization | 1 | 1 | Feb 21, 2026 | |
| VLM-as-a-Judge | 1 | 1 | Apr 21, 2026 |
| Synthetic token manipulation | 1 | 1 | Feb 21, 2026 |
|---|
| Interactive step-by-step task guidance | 1 | 1 | Feb 18, 2026 |
|---|
| Post-training Performance Evaluation | 1 | 1 | Feb 21, 2026 |
|---|
| Mathematical reasoning and calculation | 1 | 1 | Apr 21, 2026 |
|---|
| Commonsense & Factual Reasoning | 1 | 1 | Feb 27, 2026 |
|---|
| Helpful Assistant Alignment | 1 | 1 | Feb 21, 2026 |
|---|
| Perplexity Prediction | 1 | 1 | Apr 21, 2026 |
|---|
| Reddit Summary Alignment | 1 | 1 | Feb 21, 2026 |
|---|
| LLM Verification | 1 | 1 | Apr 21, 2026 |
|---|
| Reasoning and Decision-making | 1 | 1 | Feb 18, 2026 |
|---|
| Logical and Commonsense Reasoning | 1 | 1 | Feb 21, 2026 |
|---|
| Literal-to-Figurative Steering | 1 | 1 | Apr 21, 2026 |
|---|
| Goal-oriented Dialog Generation | 1 | 1 | Feb 18, 2026 |
|---|
| Instruction Following and Agent Capabilities | 1 | 1 | Feb 21, 2026 |
|---|
| Figurative-to-Literal Steering | 1 | 1 | Apr 21, 2026 |
|---|
| Conversational Parameter Extraction and Alignment | 1 | 1 | Feb 21, 2026 |
|---|
| Social Support Dialogue Generation | 1 | 1 | Apr 21, 2026 |
|---|