Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Math Word Problem Preference | 2 | 1 | Apr 14, 2026 | |
| Broad Knowledge Question Answering | 2 | 1 | Apr 14, 2026 | |
| Predicting human judges' overall quality (Q0) | 2 | 1 | Feb 18, 2026 | |
| Toxic Degeneration | 2 | 1 | Feb 18, 2026 | |
| Role-playing performance evaluation | 2 | 1 | Feb 19, 2026 | |
| User Preference | 2 | 2 | Jul 2, 2026 | |
| Aggregate Reasoning Evaluation | 2 | 2 | May 26, 2026 | |
| Context Traceback | 2 | 1 | Apr 14, 2026 | |
| Dialog Naturalness |
| 2 |
| 1 |
| Feb 19, 2026 |
| Instruction Tuning Evaluation | 2 | 2 | Apr 9, 2026 |
|---|
| User Simulation Quality Assessment | 2 | 1 | Feb 19, 2026 |
|---|
| Long-context Multi-modal Understanding | 2 | 2 | May 14, 2026 |
|---|
| Language-conditioned simulation | 2 | 1 | Apr 14, 2026 |
|---|
| Rule Following | 2 | 1 | Apr 13, 2026 |
|---|
| Addition reasoning | 2 | 1 | Feb 18, 2026 |
|---|
| Multi-modal Instruction Following | 2 | 4 | Apr 7, 2026 |
|---|
| Consultation Capability Evaluation | 2 | 1 | Apr 13, 2026 |
|---|
| Language Conditioned Transfer | 2 | 1 | Apr 10, 2026 |
|---|
| Open Domain | 2 | 4 | May 14, 2026 |
|---|
| Maximum reasoning | 2 | 1 | Feb 18, 2026 |
|---|
| LLM Serving Efficiency | 2 | 1 | Apr 10, 2026 |
|---|
| Multi-agent contract design | 2 | 1 | Apr 10, 2026 |
|---|
| Single-agent contract design | 2 | 1 | Apr 10, 2026 |
|---|
| Logical Reasoning and Reading Comprehension | 2 | 1 | Apr 14, 2026 |
|---|
| Memory Updating | 2 | 3 | Jun 25, 2026 |
|---|