Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Multilingual Commonsense Reasoning | 2 | 2 | May 19, 2026 | |
| Reward-wise QA fairness and alignment | 2 | 1 | Apr 7, 2026 | |
| Ordinal Preference Alignment | 2 | 1 | Apr 7, 2026 | |
| Chit-chat conversation evaluation correlation | 2 | 1 | Apr 7, 2026 | |
| Reference-free Conversation Evaluation | 2 | 1 | Apr 7, 2026 | |
| Multi-hop Retrieval-Augmented Generation | 2 | 1 | Apr 7, 2026 | |
| Human Preferences | 2 | 1 | Feb 18, 2026 | |
| Generalization to Unseen Preferences | 2 | 1 | Apr 7, 2026 | |
| Expert Preference Pairwise |
| 2 |
| 1 |
| Apr 7, 2026 |
| Role-Play Evaluation | 2 | 5 | Jun 26, 2026 |
|---|
| Role-Play | 2 | 3 | Jun 15, 2026 |
|---|
| Friend Recommendation | 2 | 1 | Apr 7, 2026 |
|---|
| Rational Ordering Evaluation | 2 | 1 | Feb 21, 2026 |
|---|
| Coherence evaluation | 2 | 2 | Mar 4, 2026 |
|---|
| LLM steering evaluation | 2 | 1 | Apr 7, 2026 |
|---|
| Dialogue Annotation | 2 | 1 | Apr 6, 2026 |
|---|
| Reward Model Controllability | 2 | 1 | Apr 7, 2026 |
|---|
| Language Understanding and Question Answering | 2 | 2 | Jul 2, 2026 |
|---|
| Multi-Agent Collaboration | 2 | 2 | Jun 12, 2026 |
|---|
| Next Step Generation | 2 | 3 | May 15, 2026 |
|---|
| Adversarial Toxicity Refusal | 2 | 1 | Apr 3, 2026 |
|---|
| General Reasoning Evaluation | 2 | 2 | Apr 7, 2026 |
|---|
| Policy Alignment | 2 | 1 | Feb 21, 2026 |
|---|
| Thought Generation | 2 | 1 | Feb 18, 2026 |
|---|
| Mathematical and Knowledge Reasoning | 2 | 1 | Feb 18, 2026 |
|---|