Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| LLM Inference Efficiency | 13 | 10 | Jul 7, 2026 | |
| Interactive Video Generation | 13 | 10 | Jun 17, 2026 | |
| Proactive Assistance | 13 | 4 | Jun 4, 2026 | |
| Misalignment Detection | 13 | 4 | Jul 3, 2026 | |
| Long-context Memory Retrieval and Reasoning | 13 | 3 | Jun 11, 2026 | |
| Story Continuation | 13 | 3 | Feb 18, 2026 | |
| Downstream Task Evaluation | 13 | 14 | Jul 7, 2026 | |
| Zero-shot Question Answering | 13 | 12 | Jun 15, 2026 | |
| Model Ranking Prediction |
| 13 |
| 2 |
| May 19, 2026 |
| Dialog | 13 | 5 | May 26, 2026 |
|---|
| Mobile Device Operation | 12 | 1 | Feb 18, 2026 |
|---|
| Step-wise Verification | 12 | 2 | May 21, 2026 |
|---|
| Medical Knowledge Evaluation | 12 | 5 | May 8, 2026 |
|---|
| Next step prediction | 12 | 4 | Apr 27, 2026 |
|---|
| Scenario Generation | 12 | 5 | Jun 30, 2026 |
|---|
| Few-shot example selection | 12 | 1 | Feb 18, 2026 |
|---|
| Time to First Token (TTFT) | 12 | 4 | Jun 19, 2026 |
|---|
| LLM-as-a-judge evaluation | 12 | 6 | Jun 19, 2026 |
|---|
| Emotional Support Conversation | 12 | 11 | Apr 21, 2026 |
|---|
| Agent evaluation | 12 | 5 | Jun 24, 2026 |
|---|
| Creative Generation | 12 | 6 | Jun 30, 2026 |
|---|
| Instruction Hierarchy Robustness | 12 | 1 | Mar 12, 2026 |
|---|
| Generative Task | 12 | 4 | Jun 23, 2026 |
|---|
| Hallucination assessment | 12 | 22 | Jun 11, 2026 |
|---|
| Memory | 12 | 4 | Jun 2, 2026 |
|---|