Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Zero-shot downstream reasoning and question answering | 1 | 1 | Feb 22, 2026 | |
| Mathematical reasoning with verbally expressed problems | 1 | 1 | Apr 29, 2026 | |
| Social Agent Evaluation | 1 | 1 | Feb 22, 2026 | |
| Tool Learning under Instruction with Missing Key Information | 1 | 1 | Apr 30, 2026 | |
| Regular-Length Story Visualization | 1 | 1 | Feb 22, 2026 | |
| Tool Learning under Instruction with Multiple Requests | 1 | 1 | Apr 30, 2026 | |
| Tool Learning under Instruction with Error | 1 | 1 | Apr 30, 2026 | |
| Long Story Visualization |
| 1 |
| 1 |
| Feb 22, 2026 |
| Tool Learning under Instructions Beyond Tool Capabilities | 1 | 1 | Apr 30, 2026 |
|---|
| Grade school mathematics reasoning | 1 | 1 | Feb 18, 2026 |
|---|
| Human feedback evaluation consistency | 1 | 1 | Apr 30, 2026 |
|---|
| Generative AI evaluation consistency | 1 | 1 | Apr 30, 2026 |
|---|
| Speculative Decoding General Performance | 1 | 1 | Apr 30, 2026 |
|---|
| Multi-hop commonsense reasoning | 1 | 1 | Feb 18, 2026 |
|---|
| Task#3 | 1 | 1 | Apr 30, 2026 |
|---|
| Task#4 | 1 | 1 | Apr 30, 2026 |
|---|
| Task#6 | 1 | 1 | Apr 30, 2026 |
|---|
| Mathematical Dialogue Evaluation | 1 | 1 | Feb 22, 2026 |
|---|
| Task#7 | 1 | 1 | Apr 30, 2026 |
|---|
| Task#8 | 1 | 1 | Apr 30, 2026 |
|---|
| True-or-False | 1 | 1 | Feb 18, 2026 |
|---|
| Mathematical Dialogue | 1 | 1 | Mar 25, 2026 |
|---|
| Summary-style Queries | 1 | 1 | Apr 30, 2026 |
|---|
| Summary-style Question Answering | 1 | 1 | Apr 30, 2026 |
|---|
| Chain-of-Thought Compression | 1 | 1 | Apr 30, 2026 |
|---|