Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Zero-shot Multiple Choice Question Answering and Reasoning | 1 | 1 | Feb 19, 2026 | |
| Learning Outcome Evaluation | 1 | 1 | Apr 7, 2026 | |
| Reward Hacking Analysis | 1 | 1 | Apr 7, 2026 | |
| Constrained Reinforcement Learning for Tutoring Curricula | 1 | 1 | Apr 7, 2026 | |
| reasoning tasks | 1 | 1 | Feb 18, 2026 | |
| Model Helpfulness Evaluation | 1 | 1 | Feb 18, 2026 | |
| Generative Language Tasks | 1 | 1 | Apr 7, 2026 | |
| Generative Language Modeling and Problem Solving |
| 1 |
| 1 |
| Apr 7, 2026 |
| GUI Agent Evaluation | 1 | 1 | Apr 7, 2026 |
|---|
| Counselor Competence Assessment | 1 | 1 | Apr 7, 2026 |
|---|
| Turing Test | 1 | 1 | Feb 18, 2026 |
|---|
| Proficiency-Level Control | 1 | 1 | Apr 7, 2026 |
|---|
| Preference Controllability | 1 | 1 | Apr 7, 2026 |
|---|
| Financial Assistance Chatbot Response Generation | 1 | 1 | Feb 19, 2026 |
|---|
| HH-RLHF | 1 | 1 | Apr 7, 2026 |
|---|
| Reward Model Suitability Audit | 1 | 1 | Feb 19, 2026 |
|---|
| Helpfulness and Informativeness Assessment | 1 | 1 | Apr 7, 2026 |
|---|
| Tutor Preference Assessment | 1 | 1 | Apr 7, 2026 |
|---|
| Multimodal Intelligence | 1 | 1 | Apr 7, 2026 |
|---|
| Fine-Grained LLM-Generated Text Detection | 1 | 1 | Apr 7, 2026 |
|---|
| Multi-task Language Proficiency | 1 | 1 | Apr 8, 2026 |
|---|
| Interactive Code Assistance | 1 | 1 | Apr 8, 2026 |
|---|
| Spatial Awareness via Audio-Visual LLMs | 1 | 1 | Apr 8, 2026 |
|---|
| Memory Fidelity Evaluation | 1 | 1 | Feb 19, 2026 |
|---|
| Social Commonsense Question Answering | 1 | 1 | Apr 8, 2026 |
|---|