Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Reward Scoring | 2 | 2 | Apr 6, 2026 | |
| Alignment Reward Evaluation | 2 | 1 | Mar 24, 2026 | |
| Standard Operating Procedure execution | 2 | 1 | Mar 25, 2026 | |
| High-level instruction execution | 2 | 1 | Mar 24, 2026 | |
| Generative multiple-choice | 2 | 1 | Feb 18, 2026 | |
| Long-form reasoning | 2 | 1 | Feb 18, 2026 | |
| Persona Consistency | 2 | 2 | Mar 27, 2026 | |
| Behavior Generation | 2 | 1 | Mar 9, 2026 | |
| Confidence Alignment |
| 2 |
| 2 |
| May 13, 2026 |
| Helpfulness Assessment | 2 | 2 | May 8, 2026 |
|---|
| Instruction State Tracking | 2 | 1 | Feb 18, 2026 |
|---|
| Counseling Dialogue Evaluation | 2 | 2 | Apr 23, 2026 |
|---|
| Large Language Model Debiasing | 2 | 1 | Mar 20, 2026 |
|---|
| General Downstream Evaluation | 2 | 2 | Jun 16, 2026 |
|---|
| Personalized Review Writing | 2 | 1 | Mar 10, 2026 |
|---|
| Prompt Hygiene Evaluation | 2 | 1 | Mar 23, 2026 |
|---|
| LLM-judge evaluation | 2 | 2 | Mar 16, 2026 |
|---|
| Open-domain dialogue red teaming | 2 | 1 | Feb 18, 2026 |
|---|
| Shot-Language Understanding | 2 | 1 | Mar 20, 2026 |
|---|
| Audio Instruction Following | 2 | 2 | May 20, 2026 |
|---|
| Context Management | 2 | 2 | Jul 2, 2026 |
|---|
| Conflict Measurement | 2 | 1 | Mar 20, 2026 |
|---|
| Zero-shot Reasoning and Knowledge | 2 | 2 | Apr 14, 2026 |
|---|
| Prompt continuation | 2 | 1 | Feb 18, 2026 |
|---|
| LLM Agent Reasoning | 2 | 1 | Mar 10, 2026 |
|---|