Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Zero-shot cross-task generalization | 1 | 1 | Feb 18, 2026 | |
| Single Product shopping assistance | 1 | 1 | Mar 17, 2026 | |
| Add-on Deals shopping assistance | 1 | 1 | Mar 17, 2026 | |
| Empathetic Dialogue Evaluation | 1 | 1 | Feb 19, 2026 | |
| Future Event Prediction | 1 | 1 | Mar 17, 2026 | |
| Needle In A Haystack (NIAH) video scenario | 1 | 1 | Mar 17, 2026 | |
| Long-context Information Retrieval | 1 | 1 | Feb 18, 2026 | |
| LLM Evaluation Efficiency | 1 | 1 |
| Feb 19, 2026 |
| Reasoning Performance Aggregation | 1 | 1 | Mar 17, 2026 |
|---|
| Zero-shot Downstream Reasoning | 1 | 1 | Mar 17, 2026 |
|---|
| Assistant Response Alignment (Helpfulness and Harmlessness) | 1 | 5 | Feb 21, 2026 |
|---|
| Alignment Quality | 1 | 1 | Feb 18, 2026 |
|---|
| User trait simulation in agentic models | 1 | 1 | Mar 18, 2026 |
|---|
| Mathematical function reasoning | 1 | 1 | Feb 18, 2026 |
|---|
| Multitask Thai Language Evaluation | 1 | 1 | Feb 19, 2026 |
|---|
| Health-domain instruction following | 1 | 1 | Mar 18, 2026 |
|---|
| Long-context Temporal Reasoning | 1 | 1 | Feb 19, 2026 |
|---|
| Long-form research report generation | 1 | 1 | Mar 18, 2026 |
|---|
| completion task | 1 | 1 | Feb 18, 2026 |
|---|
| Long-form deep research | 1 | 1 | Feb 19, 2026 |
|---|
| Multimodal Large Language Model Safety Evaluation | 1 | 1 | Mar 18, 2026 |
|---|
| Output Consistency | 1 | 1 | Mar 18, 2026 |
|---|
| Multi-attribute controllable dialogue generation | 1 | 1 | Feb 18, 2026 |
|---|
| Multimodal Dialogue Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Instruction-tuned Machine Translation | 1 | 1 | Mar 18, 2026 |
|---|