Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Language Model Inference Efficiency | 1 | 1 | Mar 4, 2026 | |
| Puzzle | 1 | 1 | May 12, 2026 | |
| Sycophancy-Induced Spiral Dynamics Intervention | 1 | 1 | May 12, 2026 | |
| Pointwise Grading | 1 | 1 | Feb 18, 2026 | |
| Vision-Language Conversation | 1 | 1 | Mar 4, 2026 | |
| Multi-hop verification | 1 | 1 | May 12, 2026 | |
| LLM Serving and Question Answering | 1 | 1 | May 12, 2026 | |
| Single-Needle-in-a-Haystack | 1 | 1 | May 12, 2026 | |
| Fine-tuning Robustness against Harmful Data Attacks |
| 1 |
| 1 |
| Mar 4, 2026 |
| Snowflake Sudoku Solving | 1 | 1 | May 12, 2026 |
|---|
| Quantization-Aware Training Efficiency Analysis | 1 | 1 | May 12, 2026 |
|---|
| Multi-subject Knowledge Reasoning | 1 | 1 | May 12, 2026 |
|---|
| Judgment | 1 | 1 | Feb 18, 2026 |
|---|
| Unlearning Forgetting and Self-awareness | 1 | 1 | May 12, 2026 |
|---|
| Retention Utility | 1 | 1 | May 12, 2026 |
|---|
| User-Centric Agent Interaction | 1 | 1 | Mar 4, 2026 |
|---|
| Multi-turn Stability and Self-expression | 1 | 1 | May 12, 2026 |
|---|
| Long-Horizon User-Centric Interaction | 1 | 1 | Mar 4, 2026 |
|---|
| Multi-Turn Jailbreak Robustness | 1 | 1 | May 12, 2026 |
|---|
| Language Model Robustness | 1 | 1 | Feb 27, 2026 |
|---|
| Long-form response generation for clinical questions | 1 | 1 | Feb 18, 2026 |
|---|
| Agentic Text Synthesis | 1 | 1 | Mar 4, 2026 |
|---|
| LLM Jailbreaking Attack | 1 | 1 | May 12, 2026 |
|---|
| Self-optimization | 1 | 1 | May 12, 2026 |
|---|
| Judge Performance Evaluation | 1 | 1 | Mar 4, 2026 |
|---|