Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Fermi Problem Solving | 1 | 1 | Feb 18, 2026 | |
| Multitask Evaluation | 1 | 1 | Apr 23, 2026 | |
| Medical AI Performance Evaluation | 1 | 1 | Apr 23, 2026 | |
| General Assistant Task Solving | 1 | 1 | Feb 21, 2026 | |
| Autonomous Agent Problem Solving | 1 | 1 | Apr 23, 2026 | |
| Patient Simulation Fidelity | 1 | 1 | Feb 21, 2026 | |
| Multi-hop Reasoning and Fact-checking | 1 | 1 | Apr 23, 2026 | |
| Simulated decision-making | 1 | 1 | Apr 23, 2026 | |
| Multi-turn Dialogue Quality |
|---|
| 1 |
| 1 |
| Feb 21, 2026 |
| Zero-shot Language Reasoning | 1 | 1 | Apr 23, 2026 |
|---|
| Truthful QA | 1 | 6 | Feb 18, 2026 |
|---|
| Dialog Critique and Refinement | 1 | 1 | Feb 18, 2026 |
|---|
| Intent-Based Interaction Generation | 1 | 1 | Feb 18, 2026 |
|---|
| multi-turn interaction-based problem solving | 1 | 1 | Feb 18, 2026 |
|---|
| General Reasoning Generalization | 1 | 1 | Feb 18, 2026 |
|---|
| Agent Communication Language Evaluation | 1 | 1 | Feb 21, 2026 |
|---|
| Social Deduction Game | 1 | 1 | Apr 23, 2026 |
|---|
| Pedagogical Dialogue Classification | 1 | 1 | Feb 21, 2026 |
|---|
| Multi-edit Instruction Adherence | 1 | 1 | Apr 23, 2026 |
|---|
| Autonomous LLM Agent Verification | 1 | 1 | Feb 21, 2026 |
|---|
| Collaborative decision-making | 1 | 1 | Apr 23, 2026 |
|---|
| Ablation hypothesis quality evaluation | 1 | 1 | Apr 23, 2026 |
|---|
| Malicious Requestor Interaction Selection | 1 | 1 | Feb 21, 2026 |
|---|
| Collaborative text generation | 1 | 1 | Apr 23, 2026 |
|---|
| Interactive Mathematical Reasoning | 1 | 1 | Apr 23, 2026 |
|---|