Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Conversational Intervention | 1 | 1 | May 8, 2026 | |
| Long-term Reasoning | 1 | 1 | Feb 18, 2026 | |
| Multi-Task Instruct-Tuning | 1 | 1 | May 8, 2026 | |
| LLM Agent Safety | 1 | 1 | May 8, 2026 | |
| Chat Model Evaluation | 1 | 1 | May 8, 2026 | |
| Commonsense and Mathematical Reasoning | 1 | 1 | Feb 24, 2026 | |
| Logical and Mathematical Reasoning under Counterfactuals | 1 | 1 | Feb 18, 2026 | |
| Suffix completion under prefix compression | 1 | 1 | May 8, 2026 | |
| RLHF Safety Evaluation |
| 1 |
| 1 |
| May 8, 2026 |
| Text-based embodied task completion | 1 | 1 | May 8, 2026 |
|---|
| Dialogue Quality Assessment | 1 | 1 | Feb 18, 2026 |
|---|
| Dialogue Alignment Evaluation | 1 | 1 | Feb 24, 2026 |
|---|
| Stage-aware Prefill | 1 | 1 | May 8, 2026 |
|---|
| Prefill KV-cache memory measurement | 1 | 1 | May 8, 2026 |
|---|
| Listwise Judging | 1 | 1 | Feb 24, 2026 |
|---|
| Sequential-insert knowledge insertion | 1 | 1 | May 8, 2026 |
|---|
| orthography_starts_with | 1 | 1 | Feb 18, 2026 |
|---|
| LLM Judge Policy Invariance Evaluation | 1 | 1 | May 8, 2026 |
|---|
| Forgetting-aware Instruction Tuning | 1 | 1 | May 8, 2026 |
|---|
| Code-Specific Instruction Tuning Evaluation | 1 | 1 | May 8, 2026 |
|---|
| Reasoning 1-speaker | 1 | 1 | Feb 24, 2026 |
|---|
| High-Level Expert Knowledge Evaluation | 1 | 1 | May 8, 2026 |
|---|
| General Instruction Tuning | 1 | 1 | May 8, 2026 |
|---|
| Reasoning 2-speaker | 1 | 1 | Feb 24, 2026 |
|---|
| Action Reasoning | 1 | 1 | May 8, 2026 |
|---|