Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Persona Simulation Naturalness Evaluation | 1 | 1 | Mar 4, 2026 | |
| Prospective Reasoning | 1 | 1 | May 13, 2026 | |
| Persona Adherence Alignment | 1 | 1 | Mar 4, 2026 | |
| LLM Calibration | 1 | 1 | Feb 18, 2026 | |
| Prompt Alignment | 1 | 1 | Mar 4, 2026 | |
| Safety Behavior Evaluation | 1 | 1 | May 13, 2026 | |
| Multi-character story generation | 1 | 1 | May 13, 2026 | |
| Honesty-Helpfulness Alignment Evaluation | 1 | 1 | Feb 18, 2026 | |
| Multi-turn omni-modal dialog |
| 1 |
| 1 |
| Mar 4, 2026 |
| Interactive Persuasive Dialogue | 1 | 1 | Feb 18, 2026 |
|---|
| Long-term Agent Memory Evaluation | 1 | 1 | May 13, 2026 |
|---|
| Memory QA | 1 | 1 | May 13, 2026 |
|---|
| Incident Management | 1 | 1 | May 13, 2026 |
|---|
| Aggregate Quality Recovery | 1 | 1 | May 13, 2026 |
|---|
| Safe RLHF Alignment | 1 | 1 | Mar 5, 2026 |
|---|
| Long-conversation Question Answering | 1 | 1 | May 13, 2026 |
|---|
| Refusal mechanism ablation | 1 | 1 | May 13, 2026 |
|---|
| Role-playing Multiple-choice Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Safe and Helpful Response Generation | 1 | 1 | Mar 5, 2026 |
|---|
| Interleaved Multi-Image Instruction Following | 1 | 1 | May 13, 2026 |
|---|
| Human Evaluation of Multilingual Capabilities | 1 | 1 | Feb 18, 2026 |
|---|
| Memory Agent Performance | 1 | 3 | Jun 11, 2026 |
|---|
| Safety and Helpfulness | 1 | 1 | May 13, 2026 |
|---|
| LLM Monitorability | 1 | 1 | May 13, 2026 |
|---|
| Ground Truth Alignment and Conversationality | 1 | 1 | Mar 5, 2026 |
|---|