Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Agentic medical interaction | 2 | 1 | Apr 8, 2026 | |
| General Reasoning Average | 2 | 2 | Jun 2, 2026 | |
| Proof writing | 2 | 2 | Jun 26, 2026 | |
| Multi-Session Dialogue Generation | 2 | 2 | Jun 11, 2026 | |
| Generalization to Unseen Preferences | 2 | 1 | Apr 7, 2026 | |
| Reward Model Controllability | 2 | 1 | Apr 7, 2026 | |
| Best-of-N Selection | 2 | 1 | Feb 20, 2026 | |
| Math Question Answering | 2 | 2 | Mar 10, 2026 | |
| Cross-Lingual Knowledge Alignment |
| 2 |
| 1 |
| Feb 18, 2026 |
| Preference Discrimination | 2 | 1 | Feb 20, 2026 |
|---|
| Expert Preference Pairwise | 2 | 1 | Apr 7, 2026 |
|---|
| Risk evaluation in caregiving responses | 2 | 1 | Feb 20, 2026 |
|---|
| Chit-chat conversation evaluation correlation | 2 | 1 | Apr 7, 2026 |
|---|
| Ordinal Preference Alignment | 2 | 1 | Apr 7, 2026 |
|---|
| Reference-free Conversation Evaluation | 2 | 1 | Apr 7, 2026 |
|---|
| Human Pairwise Comparison | 2 | 1 | Feb 20, 2026 |
|---|
| Reward-wise QA fairness and alignment | 2 | 1 | Apr 7, 2026 |
|---|
| Multi-hop Retrieval-Augmented Generation | 2 | 1 | Apr 7, 2026 |
|---|
| Friend Recommendation | 2 | 1 | Apr 7, 2026 |
|---|
| Price Negotiation | 2 | 1 | Apr 14, 2026 |
|---|
| Scientific Verification | 2 | 1 | Apr 8, 2026 |
|---|
| LLM steering evaluation | 2 | 1 | Apr 7, 2026 |
|---|
| Language Understanding and Question Answering | 2 | 2 | Jul 2, 2026 |
|---|
| Evaluating Context Influence and Input Regurgitation | 2 | 1 | Feb 27, 2026 |
|---|
| Dialogue Annotation | 2 | 1 | Apr 6, 2026 |
|---|