Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| General Reasoning Evaluation | 2 | 2 | Apr 7, 2026 | |
| Policy Alignment | 2 | 1 | Feb 21, 2026 | |
| Thought Generation | 2 | 1 | Feb 18, 2026 | |
| Mathematical and Knowledge Reasoning | 2 | 1 | Feb 18, 2026 | |
| Math & Knowledge | 2 | 3 | Feb 19, 2026 | |
| Multi-hop Retrieval-Augmented Generation | 2 | 1 | Apr 7, 2026 | |
| Friend Recommendation | 2 | 1 | Apr 7, 2026 | |
| Model Evaluation | 2 | 2 | Feb 22, 2026 | |
| Chit-chat conversation evaluation correlation |
| 2 |
| 1 |
| Apr 7, 2026 |
| Language Understanding and Question Answering | 2 | 2 | Jul 2, 2026 |
|---|
| LLM steering evaluation | 2 | 1 | Apr 7, 2026 |
|---|
| Dialogue Annotation | 2 | 1 | Apr 6, 2026 |
|---|
| Response Evaluation | 2 | 2 | Apr 23, 2026 |
|---|
| Multi-Agent Collaboration | 2 | 2 | Jun 12, 2026 |
|---|
| Accuracy Evaluation | 2 | 1 | Feb 21, 2026 |
|---|
| Verifiable Instruction Following | 2 | 2 | May 29, 2026 |
|---|
| Open-domain Dialogue Evaluation | 2 | 1 | Feb 18, 2026 |
|---|
| Long-context language evaluation | 2 | 3 | May 25, 2026 |
|---|
| Consistency Analysis | 2 | 2 | Feb 18, 2026 |
|---|
| Prompt Selection | 2 | 2 | Mar 23, 2026 |
|---|
| Adversarial Toxicity Refusal | 2 | 1 | Apr 3, 2026 |
|---|
| Correctness Assessment | 2 | 1 | Apr 2, 2026 |
|---|
| Instruction-Guided Grading | 2 | 1 | Apr 2, 2026 |
|---|
| Agent Planning and API Calling | 2 | 1 | Apr 2, 2026 |
|---|
| Static Multi-Session QA | 2 | 1 | Apr 2, 2026 |
|---|