Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| General Reasoning Average | 2 | 2 | Jun 2, 2026 | |
| Chit-chat conversation evaluation correlation | 2 | 1 | Apr 7, 2026 | |
| Expert Preference Pairwise | 2 | 1 | Apr 7, 2026 | |
| Reference-free Conversation Evaluation | 2 | 1 | Apr 7, 2026 | |
| Multi-hop Retrieval-Augmented Generation | 2 | 1 | Apr 7, 2026 | |
| Language Understanding and Question Answering | 2 | 2 | Jul 2, 2026 | |
| LLM steering evaluation | 2 | 1 | Apr 7, 2026 | |
| Dialogue Annotation | 2 | 1 | Apr 6, 2026 | |
| Game Solving |
| 2 |
| 2 |
| May 12, 2026 |
| Friend Recommendation | 2 | 1 | Apr 7, 2026 |
|---|
| Multi-Agent Collaboration | 2 | 2 | Jun 12, 2026 |
|---|
| Adversarial Toxicity Refusal | 2 | 1 | Apr 3, 2026 |
|---|
| Question Answering with Clarification | 2 | 1 | Feb 21, 2026 |
|---|
| Mathematical Reasoning Process Evaluation | 2 | 3 | May 20, 2026 |
|---|
| Crisis response generation | 2 | 1 | Apr 3, 2026 |
|---|
| Instruction-Guided Grading | 2 | 1 | Apr 2, 2026 |
|---|
| Hallucination-oriented Video Understanding | 2 | 1 | Apr 3, 2026 |
|---|
| Correctness Assessment | 2 | 1 | Apr 2, 2026 |
|---|
| Agent Planning and API Calling | 2 | 1 | Apr 2, 2026 |
|---|
| Context Traceback | 2 | 1 | Apr 14, 2026 |
|---|
| Short-form open-domain QA | 2 | 1 | Apr 2, 2026 |
|---|
| Offline Constrained RLHF | 2 | 1 | Apr 2, 2026 |
|---|
| Static Multi-Session QA | 2 | 1 | Apr 2, 2026 |
|---|
| Automated Peer Review | 2 | 1 | Feb 21, 2026 |
|---|
| Knowledge & QA | 2 | 2 | Jun 11, 2026 |
|---|