Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Language model evaluation across subsets | 1 | 1 | Feb 19, 2026 | |
| Long-context dialogue evaluation | 1 | 1 | Mar 25, 2026 | |
| Instruction Satisfaction | 1 | 1 | Feb 18, 2026 | |
| Code-related memory dialogue | 1 | 1 | Mar 25, 2026 | |
| General multi-turn dialogue evaluation | 1 | 1 | Mar 25, 2026 | |
| Persona-based memory dialogue | 1 | 1 | Mar 25, 2026 | |
| Reasoning-based Question Answering | 1 | 1 | Mar 25, 2026 | |
| Aggregated Reasoning Evaluation | 1 | 1 | Mar 25, 2026 | |
| Rational Agreement |
| 1 |
| 1 |
| Mar 25, 2026 |
| General MLLM Evaluation | 1 | 1 | Mar 25, 2026 |
|---|
| Vision-Language Conversation and Reasoning | 1 | 1 | Mar 25, 2026 |
|---|
| Text-based reasoning | 1 | 1 | Mar 26, 2026 |
|---|
| Knowledge & Understanding | 1 | 1 | Feb 19, 2026 |
|---|
| General Knowledge and Language Understanding | 1 | 2 | Apr 2, 2026 |
|---|
| Korean LLM Evaluation | 1 | 1 | Feb 19, 2026 |
|---|
| Evaluation Criterion Generation | 1 | 1 | Mar 26, 2026 |
|---|
| Multi-task Complex Understanding | 1 | 1 | Feb 18, 2026 |
|---|
| LLM-as-a-Judge Robustness to Adversarial Attacks | 1 | 1 | Feb 19, 2026 |
|---|
| Document Writing | 1 | 1 | Mar 26, 2026 |
|---|
| Language Model Alignment | 1 | 1 | Mar 26, 2026 |
|---|
| Federated Fine-tuning | 1 | 1 | Mar 26, 2026 |
|---|
| Score-based Alignment | 1 | 1 | Feb 19, 2026 |
|---|
| Reasoning and Knowledge Evaluation | 1 | 1 | Mar 26, 2026 |
|---|
| Coverage-based Alignment | 1 | 1 | Feb 19, 2026 |
|---|
| Task-oriented Dialog Response Generation | 1 | 1 | Feb 18, 2026 |
|---|