Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| CBT skill scoring | 1 | 1 | Mar 20, 2026 | |
| Helpfulness, Honesty, and Harmlessness Alignment Evaluation | 1 | 1 | Feb 18, 2026 | |
| Multi-source Answer Generation | 1 | 1 | Mar 20, 2026 | |
| Dialogic Deference Evaluation | 1 | 1 | Feb 19, 2026 | |
| Long-context retrieval and synthetic reasoning | 1 | 2 | Mar 24, 2026 | |
| Large Language Model Alignment | 1 | 1 | Mar 20, 2026 | |
| Professional-level Knowledge Acquisition | 1 | 1 | Mar 20, 2026 | |
| Expert-level Understanding | 1 |
| 1 |
| Mar 20, 2026 |
| Mental Healthcare Support Response Evaluation | 1 | 1 | Feb 19, 2026 |
|---|
| Deductive Video Reasoning | 1 | 1 | Mar 20, 2026 |
|---|
| Underspecified Prompting | 1 | 1 | Mar 20, 2026 |
|---|
| Persona-driven decision-making | 1 | 1 | Feb 19, 2026 |
|---|
| Instructional Dialogue Evaluation | 1 | 1 | Mar 20, 2026 |
|---|
| Free-format Multimodal Hallucination Assessment | 1 | 1 | Feb 18, 2026 |
|---|
| Psychological Simulation Evaluation | 1 | 1 | Feb 19, 2026 |
|---|
| Visual Sudoku | 1 | 1 | Mar 20, 2026 |
|---|
| LLM Downstream Evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| MI dialogue generation | 1 | 1 | Feb 19, 2026 |
|---|
| Function Addition | 1 | 1 | Mar 20, 2026 |
|---|
| Long-context multimodal evaluation | 1 | 1 | Feb 18, 2026 |
|---|
| Decision-Making Bias Evaluation | 1 | 1 | Mar 20, 2026 |
|---|
| General Science Reasoning | 1 | 1 | Mar 20, 2026 |
|---|
| Context-faithful Multi-hop Reasoning | 1 | 1 | Feb 18, 2026 |
|---|
| Aggregate Reasoning | 1 | 1 | Mar 20, 2026 |
|---|
| Multi-Turn User Simulation | 1 | 1 | Mar 20, 2026 |
|---|