Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Continual Multimodal Instruction Tuning | 1 | 1 | Feb 21, 2026 | |
| Harmful prompt classification | 1 | 1 | Apr 23, 2026 | |
| Multi-hop reasoning alignment | 1 | 1 | Apr 23, 2026 | |
| Status Querying | 1 | 1 | Feb 18, 2026 | |
| Multimodal app-use reasoning | 1 | 1 | Feb 21, 2026 | |
| Multimodal Large Language Model Inference | 1 | 1 | Apr 23, 2026 | |
| Alignment Robustness Evaluation | 1 | 1 | Feb 21, 2026 | |
| Tutor Leakage Evaluation |
| 1 |
| 1 |
| Apr 23, 2026 |
| Leakage Analysis in LLM-based Tutoring | 1 | 1 | Apr 23, 2026 |
|---|
| Tutor Robustness | 1 | 1 | Apr 23, 2026 |
|---|
| AI Reasoning | 1 | 1 | Apr 23, 2026 |
|---|
| Generation quality evaluation | 1 | 1 | Feb 21, 2026 |
|---|
| Logical Reasoning Verification | 1 | 1 | Apr 23, 2026 |
|---|
| Multi-turn Jailbreak Attack Robustness | 1 | 1 | Apr 23, 2026 |
|---|
| Controversy Controllability Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Output Sequence Length Prediction | 1 | 1 | Feb 18, 2026 |
|---|
| Decode | 1 | 1 | Feb 21, 2026 |
|---|
| Language-conditioned robot control | 1 | 1 | Apr 23, 2026 |
|---|
| Interactive Dialogue Management | 1 | 1 | Feb 21, 2026 |
|---|
| 1-shot Learning | 1 | 1 | Apr 23, 2026 |
|---|
| Buyer Negotiation | 1 | 1 | Feb 18, 2026 |
|---|
| multi-round investment game | 1 | 1 | Feb 21, 2026 |
|---|
| Reasoning and Generation | 1 | 1 | Apr 23, 2026 |
|---|
| Large Language Model Reasoning and Coding | 1 | 1 | Apr 23, 2026 |
|---|
| Zero-shot Boolean Question Answering | 1 | 1 | Apr 23, 2026 |
|---|