Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Long-form factuality evaluation | 3 | 1 | Apr 15, 2026 | |
| Multimodal In-Context Learning | 3 | 3 | Apr 16, 2026 | |
| Multi-agent Cognitive Orchestration | 3 | 1 | Apr 21, 2026 | |
| Linear Concept Accessibility and Steering | 3 | 1 | Apr 20, 2026 | |
| LoRA Adapter Transfer | 3 | 1 | Apr 15, 2026 | |
| Social Deduction Game Gameplay | 3 | 2 | May 29, 2026 | |
| Long-horizon legal inquiry | 3 | 1 | Feb 18, 2026 | |
| Value Modeling | 3 | 1 | Feb 18, 2026 | |
| 3 |
| 2 |
| Jun 2, 2026 |
| Self-confidence estimation | 3 | 1 | Feb 20, 2026 |
|---|
| Multi-turn Jailbreak | 3 | 2 | Jun 2, 2026 |
|---|
| Chat Evaluation | 3 | 4 | Jun 26, 2026 |
|---|
| Multimodal Medical Reasoning | 3 | 4 | Jul 1, 2026 |
|---|
| PRP Faithfulness Evaluation | 3 | 1 | Feb 18, 2026 |
|---|
| 20 Questions | 3 | 1 | Feb 19, 2026 |
|---|
| Multi-step Reasoning over Code Dependencies | 3 | 1 | Apr 14, 2026 |
|---|
| Hallucination Benchmark | 3 | 1 | Apr 14, 2026 |
|---|
| Incorrect Reasoning Path Detection | 3 | 1 | Apr 14, 2026 |
|---|
| Creative Writing Evaluation | 3 | 2 | Jun 4, 2026 |
|---|
| Hallucination Examination | 3 | 2 | Jun 11, 2026 |
|---|
| LLM-as-a-Judge Evaluation Consistency | 3 | 1 | Feb 18, 2026 |
|---|
| Strategic AI Persuasion (Sender) | 3 | 1 | Feb 19, 2026 |
|---|
| Strategic AI Persuasion (Receiver) | 3 | 1 | Feb 19, 2026 |
|---|
| Mathematical Reasoning (Calculator) | 3 | 1 | Feb 19, 2026 |
|---|
| UMUI Judgment Calibration | 3 | 1 | Apr 13, 2026 |
|---|