Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Alignment Reward Evaluation | 2 | 1 | Mar 24, 2026 | |
| Autoformalization and Proving | 2 | 1 | Mar 23, 2026 | |
| Prefill-stage hallucination risk detection | 2 | 1 | Mar 23, 2026 | |
| Prompt continuation | 2 | 1 | Feb 18, 2026 | |
| LLM Agent Reasoning | 2 | 1 | Mar 10, 2026 | |
| Aggregated LLM Evaluation | 2 | 2 | May 4, 2026 | |
| Generative Hallucination Evaluation | 2 | 3 | May 28, 2026 | |
| MultiModal Long-Context Understanding | 2 | 2 | Apr 21, 2026 | |
| Self-introduction generation |
| 2 |
| 1 |
| Feb 18, 2026 |
| Dialogue Quality | 2 | 2 | Jun 23, 2026 |
|---|
| Helpfulness Assessment | 2 | 2 | May 8, 2026 |
|---|
| Persona Role-Playing Faithfulness | 2 | 1 | Feb 18, 2026 |
|---|
| Math Question Verification | 2 | 1 | Mar 11, 2026 |
|---|
| CBT Conversation Generation | 2 | 1 | Feb 19, 2026 |
|---|
| Large Language Model Debiasing | 2 | 1 | Mar 20, 2026 |
|---|
| Multi-domain Knowledge and Reasoning | 2 | 3 | Apr 21, 2026 |
|---|
| Exam | 2 | 1 | Mar 11, 2026 |
|---|
| Rationale Faithfulness Evaluation | 2 | 2 | May 12, 2026 |
|---|
| Context Management | 2 | 2 | Jul 2, 2026 |
|---|
| Shot-Language Understanding | 2 | 1 | Mar 20, 2026 |
|---|
| Scientific Review Feedback Generation | 2 | 1 | Mar 11, 2026 |
|---|
| Audio Instruction Following | 2 | 2 | May 20, 2026 |
|---|
| Zero-shot Reasoning and Knowledge | 2 | 2 | Apr 14, 2026 |
|---|
| Prompt Hygiene Evaluation | 2 | 1 | Mar 23, 2026 |
|---|
| Conflict Measurement | 2 | 1 | Mar 20, 2026 |
|---|