Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| LLM Jailbreak Defense | 1 | 1 | May 4, 2026 | |
| Content Generation + Harmlessness | 1 | 1 | Feb 18, 2026 | |
| Hard Math Reasoning | 1 | 1 | Feb 18, 2026 | |
| Preference Domain Analysis | 1 | 1 | Feb 22, 2026 | |
| Context Length Estimation | 1 | 1 | Feb 22, 2026 | |
| General-purpose Language Evaluation | 1 | 1 | May 4, 2026 | |
| Long-horizon Repo Exploration | 1 | 1 | Feb 22, 2026 | |
| Safety Alignment (Jailbreak Resistance) | 1 | 1 | May 4, 2026 |
| Pairwise RAG Comparison | 1 | 1 | May 4, 2026 |
|---|
| Cross-lingual Fact-to-Text generation | 1 | 1 | Feb 18, 2026 |
|---|
| Long-horizon Chained Tasks | 1 | 1 | Feb 22, 2026 |
|---|
| Agent Adaptation | 1 | 1 | May 4, 2026 |
|---|
| Instruction-following generation | 1 | 1 | Feb 22, 2026 |
|---|
| Benchmark Subset Selection | 1 | 1 | May 4, 2026 |
|---|
| Controllable Model Distillation | 1 | 1 | Feb 22, 2026 |
|---|
| Due Diligence | 1 | 1 | May 4, 2026 |
|---|
| Over-refusal Assessment | 1 | 1 | Feb 18, 2026 |
|---|
| Adaptive Token Selection | 1 | 1 | May 6, 2026 |
|---|
| Fact verify | 1 | 1 | May 6, 2026 |
|---|
| LLM Compression | 1 | 1 | Feb 22, 2026 |
|---|
| LLM Detection Interpretability | 1 | 1 | May 6, 2026 |
|---|
| Role-playing Instruction Following | 1 | 1 | Feb 18, 2026 |
|---|
| Majority Vote | 1 | 1 | May 6, 2026 |
|---|
| List length | 1 | 1 | May 6, 2026 |
|---|
| Zero-shot word prediction | 1 | 1 | May 6, 2026 |
|---|