Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Proactive Task Management | 1 | 1 | May 27, 2026 | |
| Interleaved Reactive-Proactive Requests | 1 | 1 | May 27, 2026 | |
| Federated Value Alignment (Direct Preference Optimization) | 1 | 1 | Jul 7, 2026 | |
| Lipschitz Ball Verification | 1 | 1 | Apr 2, 2026 | |
| Long input summarization | 1 | 1 | Feb 18, 2026 | |
| Best-of-N Alignment Evaluation | 1 | 1 | Mar 16, 2026 | |
| Multimodal Reasoning Evaluation | 1 | 1 | Mar 16, 2026 | |
| Emotion Control | 1 | 1 |
| May 27, 2026 |
| Behavioral and latency evaluation under correction | 1 | 1 | Jul 7, 2026 |
|---|
| End-to-end Attention Latency | 1 | 1 | Apr 2, 2026 |
|---|
| Model Merging Performance Aggregation | 1 | 1 | Feb 18, 2026 |
|---|
| Intent-conditioned model selection | 1 | 1 | Apr 2, 2026 |
|---|
| Model Utility Maintenance | 1 | 1 | Feb 19, 2026 |
|---|
| Multi-Agent Governance | 1 | 1 | Mar 16, 2026 |
|---|
| Reward Model Bias Evaluation | 1 | 1 | May 27, 2026 |
|---|
| Binary Question Answering (Yes/No) | 1 | 1 | Mar 17, 2026 |
|---|
| Multiple Tasks | 1 | 1 | May 27, 2026 |
|---|
| science multiple-choice question answering | 1 | 1 | Mar 17, 2026 |
|---|
| Cross-task Knowledge Editing | 1 | 1 | Feb 18, 2026 |
|---|
| Few-shot Language Evaluation | 1 | 1 | Feb 19, 2026 |
|---|
| Output-refinement | 1 | 1 | May 28, 2026 |
|---|
| Long-term Dialogue Memory Management | 1 | 1 | May 28, 2026 |
|---|
| Instruction Following and Helpfulness Evaluation | 1 | 3 | Feb 18, 2026 |
|---|
| Constraint-following Instruction Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Long-term dialogue memory evaluation | 1 | 1 | May 28, 2026 |
|---|