Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Conversational Quality Evaluation | 1 | 1 | May 19, 2026 | |
| Math Proof Reward Modeling | 1 | 1 | Feb 18, 2026 | |
| Turn-level correlation with human ratings | 1 | 1 | May 19, 2026 | |
| Structured Commitment Generation | 1 | 1 | May 19, 2026 | |
| Zero-shot cross-task generalization | 1 | 1 | Feb 18, 2026 | |
| Winner Selection | 1 | 1 | May 19, 2026 | |
| Psychiatric dialogue evaluation | 1 | 1 | Mar 10, 2026 | |
| Model fidelity evaluation | 1 | 1 | May 19, 2026 | |
| End-to-End Negotiation Simulation |
| 1 |
| 1 |
| May 29, 2026 |
| Empathetic Dialogue Evaluation | 1 | 1 | Feb 19, 2026 |
|---|
| Cross-modal multi-expert orchestration | 1 | 1 | May 29, 2026 |
|---|
| Needle In A Haystack (NIAH) video scenario | 1 | 1 | Mar 17, 2026 |
|---|
| Multi-turn 3D Editing | 1 | 1 | May 19, 2026 |
|---|
| Mathematical Proof Reward Modeling | 1 | 1 | Feb 18, 2026 |
|---|
| Personalized Dialogue Evaluation | 1 | 1 | Mar 10, 2026 |
|---|
| Knowledge-aware refusal | 1 | 1 | May 19, 2026 |
|---|
| Personalized Answer Generation | 1 | 1 | Mar 10, 2026 |
|---|
| LLM-judge preference scoring | 1 | 1 | Feb 18, 2026 |
|---|
| Capability Retention | 1 | 1 | May 19, 2026 |
|---|
| Closed-book recall | 1 | 1 | May 19, 2026 |
|---|
| General-domain capability retention | 1 | 1 | May 19, 2026 |
|---|
| Multi-turn clinical response generation | 1 | 1 | May 19, 2026 |
|---|
| Question-based Generation | 1 | 1 | Mar 10, 2026 |
|---|
| Even Pairs | 1 | 1 | May 19, 2026 |
|---|
| Task Fulfillment | 1 | 1 | Mar 10, 2026 |
|---|