Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Knowledge Model Editing | 1 | 1 | Mar 24, 2026 | |
| Multi-agent discussion attack | 1 | 1 | Mar 24, 2026 | |
| Question Asking Policy Evaluation | 1 | 1 | Feb 19, 2026 | |
| Dialectal Bias Evaluation | 1 | 1 | Mar 24, 2026 | |
| Conversational Machine Comprehension | 1 | 1 | Feb 18, 2026 | |
| Mixed 20 Question | 1 | 1 | Feb 19, 2026 | |
| Dialogue Session Performance Analysis | 1 | 1 | Mar 24, 2026 | |
| General Ability | 1 | 1 | Feb 18, 2026 | |
| 1 |
| 1 |
| Feb 19, 2026 |
| Prompt Steering | 1 | 1 | Feb 19, 2026 |
|---|
| Multi-level multi-discipline evaluation | 1 | 2 | Feb 18, 2026 |
|---|
| Science and Knowledge Question Answering | 1 | 1 | Mar 24, 2026 |
|---|
| Commonsense and Language Reasoning | 1 | 1 | Mar 24, 2026 |
|---|
| Mathematical and Quantitative Reasoning | 1 | 1 | Mar 24, 2026 |
|---|
| Enterprise task completion | 1 | 1 | Mar 24, 2026 |
|---|
| Hard Reasoning Tasks | 1 | 2 | May 21, 2026 |
|---|
| Negotiation Dialogue | 1 | 1 | Mar 24, 2026 |
|---|
| Chinese LLM Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Long-Context Mathematical Reasoning | 1 | 1 | Mar 24, 2026 |
|---|
| NLU and Question Answering | 1 | 1 | Mar 24, 2026 |
|---|
| Algorithm Recommendation | 1 | 1 | Mar 24, 2026 |
|---|
| Memory Update | 1 | 1 | Feb 19, 2026 |
|---|
| Virtual Standardized Patient Simulation | 1 | 1 | Mar 24, 2026 |
|---|
| LLM Re-ranking | 1 | 1 | Feb 18, 2026 |
|---|
| Multiple-choice Question Reasoning | 1 | 1 | Mar 24, 2026 |
|---|