Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| General Reasoning Average | 2 | 2 | Jun 2, 2026 | |
| Open-ended QA Response Ranking | 2 | 1 | Feb 18, 2026 | |
| Ordinal Preference Alignment | 2 | 1 | Apr 7, 2026 | |
| Expert Preference Pairwise | 2 | 1 | Apr 7, 2026 | |
| Reward-wise QA fairness and alignment | 2 | 1 | Apr 7, 2026 | |
| Reference-free Conversation Evaluation | 2 | 1 | Apr 7, 2026 | |
| Preference Profile Estimation | 2 | 1 | Feb 18, 2026 | |
| Chit-chat conversation evaluation correlation | 2 | 1 | Apr 7, 2026 | |
| Multi-hop Retrieval-Augmented Generation |
| 2 |
| 1 |
| Apr 7, 2026 |
| LLM Inference Throughput | 2 | 1 | Feb 20, 2026 |
|---|
| Friend Recommendation | 2 | 1 | Apr 7, 2026 |
|---|
| Diversity measurement correlation | 2 | 1 | Feb 18, 2026 |
|---|
| LLM steering evaluation | 2 | 1 | Apr 7, 2026 |
|---|
| Chat Fine-tuning | 2 | 1 | Feb 20, 2026 |
|---|
| Language Understanding and Question Answering | 2 | 2 | Jul 2, 2026 |
|---|
| Memorization Reduction | 2 | 1 | Feb 20, 2026 |
|---|
| Synthetic Tasks | 2 | 3 | May 19, 2026 |
|---|
| Robustness against harmful content generation | 2 | 1 | Feb 20, 2026 |
|---|
| Dialogue Annotation | 2 | 1 | Apr 6, 2026 |
|---|
| Math Word Problem Reasoning | 2 | 1 | Feb 18, 2026 |
|---|
| Long-generation reasoning | 2 | 1 | Feb 20, 2026 |
|---|
| Long-context performance evaluation | 2 | 4 | Jul 2, 2026 |
|---|
| General Capability Retention | 2 | 2 | May 19, 2026 |
|---|
| Hallucination Annotation | 2 | 2 | Feb 18, 2026 |
|---|
| Multilingual Commonsense Reasoning | 2 | 2 | May 19, 2026 |
|---|