Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Preference-conditioned Autonomous Highway Navigation | 1 | 1 | Jun 18, 2026 | |
| Text-based reasoning | 1 | 1 | Mar 26, 2026 | |
| Multi-skill Compositional Reasoning | 1 | 1 | Jun 18, 2026 | |
| Long-context language model evaluation | 1 | 2 | Apr 13, 2026 | |
| Large Language Model Downstream Evaluation | 1 | 1 | May 26, 2026 | |
| Multi-turn conversational quality | 1 | 1 | May 26, 2026 | |
| Writing capabilities | 1 | 1 | May 26, 2026 | |
| Korean LLM Evaluation | 1 | 1 | Feb 19, 2026 | |
| Behavioral Accuracy |
| 1 |
| 1 |
| Jun 18, 2026 |
| Service and Safety Evaluation | 1 | 1 | May 26, 2026 |
|---|
| Memorized Content Extraction | 1 | 1 | Mar 13, 2026 |
|---|
| Online Customer Support | 1 | 1 | May 26, 2026 |
|---|
| Generative Language Model Merging | 1 | 1 | Jun 18, 2026 |
|---|
| LLM Workflow Optimization | 1 | 1 | Feb 18, 2026 |
|---|
| Online Livestream Interaction | 1 | 1 | May 26, 2026 |
|---|
| Disagreement Handling | 1 | 1 | Mar 13, 2026 |
|---|
| Mathematical Reasoning Repair | 1 | 1 | May 26, 2026 |
|---|
| Routine Task Management | 1 | 1 | Mar 13, 2026 |
|---|
| General Task Performance | 1 | 1 | Feb 19, 2026 |
|---|
| Regional knowledge and conversational settings | 1 | 1 | May 26, 2026 |
|---|
| General performance assessment | 1 | 1 | May 26, 2026 |
|---|
| Reward-oriented Decoding | 1 | 1 | Feb 18, 2026 |
|---|
| Human preference evaluation for response quality | 1 | 1 | May 26, 2026 |
|---|
| Single-turn child safety evaluation | 1 | 1 | May 26, 2026 |
|---|
| Role-play dialogue comprehension | 1 | 1 | May 26, 2026 |
|---|