Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Interaction integrity monitoring | 1 | 1 | Apr 16, 2026 | |
| Personality-based Memory | 1 | 1 | Apr 16, 2026 | |
| Mathematics Question Answering | 1 | 1 | Feb 20, 2026 | |
| Personalized Multimodal Evaluation | 1 | 1 | Apr 16, 2026 | |
| Attention Operator Latency | 1 | 1 | Feb 20, 2026 | |
| Federated Multi-task Fine-tuning | 1 | 1 | Apr 16, 2026 | |
| Federated Parameter-Efficient Fine-Tuning | 1 | 1 | Apr 16, 2026 | |
| Zero-shot Downstream Reasoning and Knowledge Tasks | 1 | 1 |
| Feb 20, 2026 |
| Tool Call Repair | 1 | 1 | Apr 16, 2026 |
|---|
| Commonsense Reasoning and Knowledge Understanding | 1 | 1 | Apr 16, 2026 |
|---|
| Causal Attribution | 1 | 1 | Apr 16, 2026 |
|---|
| Human Evaluation of Dialogue Systems | 1 | 1 | Feb 18, 2026 |
|---|
| Decoding Stability | 1 | 1 | Apr 16, 2026 |
|---|
| Compositional Understanding | 1 | 1 | Apr 16, 2026 |
|---|
| Hybrid Decision Gate | 1 | 1 | Apr 16, 2026 |
|---|
| On-device Training | 1 | 1 | Apr 16, 2026 |
|---|
| Answer quality evaluation | 1 | 1 | Apr 16, 2026 |
|---|
| Utterance-level user simulation | 1 | 1 | Apr 16, 2026 |
|---|
| Personality performance evaluation | 1 | 3 | Jun 11, 2026 |
|---|
| Efficient Reasoning | 1 | 1 | Apr 16, 2026 |
|---|
| Forgetting analysis | 1 | 1 | Apr 16, 2026 |
|---|
| Conversational Instruction Following | 1 | 1 | Apr 16, 2026 |
|---|
| General Instruction Following Evaluation | 1 | 1 | Apr 16, 2026 |
|---|
| Instruction Following and Multimodal Reasoning | 1 | 1 | Feb 20, 2026 |
|---|
| Task Gradient Conflict | 1 | 1 | Apr 16, 2026 |
|---|