Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| LLM Routing | 65 | 19 | Jun 5, 2026 | |
| Zero-shot Evaluation | 63 | 80 | Jul 9, 2026 | |
| Preference Alignment | 61 | 32 | Jul 7, 2026 | |
| Generation | 56 | 32 | Jun 29, 2026 | |
| Open-ended generation | 56 | 44 | Jul 7, 2026 | |
| Multi-hop Reasoning | 54 | 44 | Jun 30, 2026 | |
| Model Editing | 53 | 20 | Jun 18, 2026 | |
| Dialogue Response Generation | 53 | 24 | Jun 15, 2026 | |
| Response Generation |
| 49 |
| 43 |
| Jun 12, 2026 |
| LLM Inference | 49 | 22 | Jun 29, 2026 |
|---|
| Conversational Question Answering | 48 | 34 | Jun 29, 2026 |
|---|
| General Knowledge | 47 | 85 | Jul 8, 2026 |
|---|
| Instruction-based Image Editing | 47 | 27 | Jun 11, 2026 |
|---|
| Model Routing | 47 | 12 | Jun 30, 2026 |
|---|
| LLM Alignment | 42 | 26 | Jun 11, 2026 |
|---|
| Creative Writing | 41 | 35 | Jun 25, 2026 |
|---|
| Factuality Evaluation | 41 | 33 | Jun 23, 2026 |
|---|
| Long-context evaluation | 38 | 44 | Jul 9, 2026 |
|---|
| Long-form Question Answering | 38 | 28 | Jul 7, 2026 |
|---|
| Complex Reasoning | 38 | 33 | Jul 7, 2026 |
|---|
| Long-context language modeling | 38 | 61 | Jul 2, 2026 |
|---|
| Math Word Problem Solving | 37 | 46 | Jul 1, 2026 |
|---|
| Jailbreak | 37 | 26 | Jun 11, 2026 |
|---|
| Zero-shot Reasoning | 37 | 51 | Jul 7, 2026 |
|---|
| Dialogue Evaluation | 37 | 20 | Jul 7, 2026 |
|---|