Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Multi-Agent Collaboration | 2 | 2 | Jun 12, 2026 | |
| Automated Peer Review | 2 | 1 | Feb 21, 2026 | |
| Hallucination-oriented Video Understanding | 2 | 1 | Apr 3, 2026 | |
| Crisis response generation | 2 | 1 | Apr 3, 2026 | |
| Adversarial Toxicity Refusal | 2 | 1 | Apr 3, 2026 | |
| Instruction-Guided Grading | 2 | 1 | Apr 2, 2026 | |
| Correctness Assessment | 2 | 1 | Apr 2, 2026 | |
| Framework Capability Comparison | 2 | 2 | Jun 11, 2026 | |
| Agent Planning and API Calling |
|---|
| 2 |
| 1 |
| Apr 2, 2026 |
| Offline Constrained RLHF | 2 | 1 | Apr 2, 2026 |
|---|
| Static Multi-Session QA | 2 | 1 | Apr 2, 2026 |
|---|
| Short-form open-domain QA | 2 | 1 | Apr 2, 2026 |
|---|
| Long-horizon dialogue | 2 | 1 | Apr 1, 2026 |
|---|
| Speech-to-speech instruction-following | 2 | 2 | Apr 21, 2026 |
|---|
| General Model Capability | 2 | 2 | Jun 9, 2026 |
|---|
| Emotion Reasoning | 2 | 1 | Feb 21, 2026 |
|---|
| Prompt Optimization Evaluation | 2 | 1 | Apr 1, 2026 |
|---|
| Free-form Question Answering | 2 | 2 | Jun 4, 2026 |
|---|
| Content Generation | 2 | 2 | Feb 18, 2026 |
|---|
| Pluralistic Reward Model Learning | 2 | 1 | Mar 31, 2026 |
|---|
| Knowledge & QA | 2 | 2 | Jun 11, 2026 |
|---|
| Professional Reasoning | 2 | 2 | May 26, 2026 |
|---|
| Universal multi-modal reasoning | 2 | 1 | Feb 21, 2026 |
|---|
| Factuality and Reasoning | 2 | 1 | Feb 18, 2026 |
|---|
| Token Generation | 2 | 2 | Apr 10, 2026 |
|---|