Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Charity Persuasion | 1 | 1 | Apr 14, 2026 | |
| Open-style response generation | 1 | 1 | Feb 18, 2026 | |
| Self-reported learning evaluation | 1 | 1 | Feb 20, 2026 | |
| Model Size Evaluation | 1 | 1 | Feb 18, 2026 | |
| Text-to-Text Creative Reasoning | 1 | 1 | Feb 18, 2026 | |
| Controllable writing | 1 | 1 | Apr 14, 2026 | |
| Hijaiyah Literacy Instruction | 1 | 1 | Feb 20, 2026 | |
| End-to-End Conversational AI | 1 | 1 | Feb 20, 2026 | |
| Persona fidelity evaluation |
| 1 |
| 1 |
| Feb 20, 2026 |
| Model-Specific Preference Bias Evaluation | 1 | 1 | Apr 14, 2026 |
|---|
| Long-sequence generative recommendation | 1 | 1 | Feb 20, 2026 |
|---|
| Dialogue-level Quality Assessment | 1 | 1 | Feb 18, 2026 |
|---|
| Adversarial Theory of Mind | 1 | 1 | Apr 14, 2026 |
|---|
| Conversation Performance | 1 | 1 | Feb 20, 2026 |
|---|
| Interactive Social Privacy and Theory of Mind | 1 | 1 | Apr 14, 2026 |
|---|
| Murder Mystery Game Evaluation | 1 | 1 | Apr 14, 2026 |
|---|
| Multi-turn conversational dataset recommendation | 1 | 1 | Feb 20, 2026 |
|---|
| Empathic Dialogue | 1 | 1 | Apr 14, 2026 |
|---|
| Knowledge Gap Identification | 1 | 1 | Apr 14, 2026 |
|---|
| Individualistic human value reasoning | 1 | 1 | Feb 18, 2026 |
|---|
| Reasoning and Persona Consistency | 1 | 1 | Feb 20, 2026 |
|---|
| Epistemic Quality Evaluation | 1 | 1 | Apr 14, 2026 |
|---|
| HLE | 1 | 1 | Apr 14, 2026 |
|---|
| Research Evaluation | 1 | 1 | Apr 14, 2026 |
|---|
| Downstream task performance correlation | 1 | 1 | Feb 18, 2026 |
|---|