Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Helpfulness Assessment | 2 | 2 | May 8, 2026 | |
| Multi-discipline Understanding | 2 | 4 | May 14, 2026 | |
| Web-based Reasoning | 2 | 2 | Jun 18, 2026 | |
| Social Norm Dialogue Generation | 2 | 1 | Mar 13, 2026 | |
| Rationale Faithfulness Evaluation | 2 | 2 | May 12, 2026 | |
| Conflict Measurement | 2 | 1 | Mar 20, 2026 | |
| Audio Instruction Following | 2 | 2 | May 20, 2026 | |
| Knowledge-focused evaluation | 2 | 1 | Feb 19, 2026 | |
| Context Management |
| 2 |
| 2 |
| Jul 2, 2026 |
| Zero-shot Reasoning and Knowledge | 2 | 2 | Apr 14, 2026 |
|---|
| Missing Premise Test | 2 | 1 | Feb 19, 2026 |
|---|
| High-level instruction execution | 2 | 1 | Mar 24, 2026 |
|---|
| Large Language Model Debiasing | 2 | 1 | Mar 20, 2026 |
|---|
| Shot-Language Understanding | 2 | 1 | Mar 20, 2026 |
|---|
| Safety Dialogue Evaluation | 2 | 2 | Mar 25, 2026 |
|---|
| MoE LLM Serving | 2 | 1 | Jun 23, 2026 |
|---|
| Average across tasks | 2 | 1 | Mar 19, 2026 |
|---|
| Downstream Performance Evaluation | 2 | 4 | May 22, 2026 |
|---|
| Long-CoT Reasoning | 2 | 1 | Mar 13, 2026 |
|---|
| Scientific Research Automation | 2 | 1 | Mar 19, 2026 |
|---|
| Knowledge-enhanced Text Generation | 2 | 1 | Mar 18, 2026 |
|---|
| Biomedical domain adaptation | 2 | 1 | Mar 16, 2026 |
|---|
| Personalized Story Generation | 2 | 1 | Mar 18, 2026 |
|---|
| Multi-turn financial advisory dialogue | 2 | 2 | Jun 30, 2026 |
|---|
| Big-Bench Hard | 2 | 2 | Apr 21, 2026 |
|---|