Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Open-ended Visual Chat | 1 | 1 | Feb 18, 2026 | |
| Conversational Agent Interaction | 1 | 1 | Mar 23, 2026 | |
| Length Following | 1 | 1 | Mar 23, 2026 | |
| Long-generation | 1 | 1 | Mar 23, 2026 | |
| Short-length Instruction Following | 1 | 1 | Mar 23, 2026 | |
| Psychotherapy Dialogue Evaluation | 1 | 1 | Feb 18, 2026 | |
| Clinical Medical Reasoning Evaluation | 1 | 1 | Feb 19, 2026 | |
| Zero-shot Task Completion | 1 | 1 | Mar 23, 2026 | |
| Language Instruction Following |
| 1 |
| 1 |
| Feb 18, 2026 |
| Conversation2Note | 1 | 1 | Feb 18, 2026 |
|---|
| Overall Language Performance | 1 | 2 | Feb 19, 2026 |
|---|
| Zero-shot long-context reasoning | 1 | 1 | Mar 23, 2026 |
|---|
| Synthetic Dialogue Generation Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Factual Accuracy and Reasoning | 1 | 2 | May 26, 2026 |
|---|
| Long-context reasoning (Pairs) | 1 | 1 | Mar 23, 2026 |
|---|
| Client Simulation Receptivity Consistency | 1 | 1 | Feb 18, 2026 |
|---|
| Overall reasoning performance | 1 | 1 | Feb 19, 2026 |
|---|
| General Language Model Capability | 1 | 1 | Mar 24, 2026 |
|---|
| RLHF Alignment Evaluation | 1 | 1 | Mar 24, 2026 |
|---|
| High-level instruction following | 1 | 1 | Mar 24, 2026 |
|---|
| College-level Multimodal Understanding | 1 | 1 | Feb 18, 2026 |
|---|
| Emotional Support Conversation Seeker Simulation | 1 | 1 | Feb 19, 2026 |
|---|
| Multi-agent aggregation | 1 | 1 | Mar 24, 2026 |
|---|
| Dialogue Diversity Evaluation | 1 | 1 | Feb 19, 2026 |
|---|
| Grounded Chess Reasoning | 1 | 1 | Mar 24, 2026 |
|---|