Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Autoformalization and Proving | 2 | 1 | Mar 23, 2026 | |
| Alignment Reward Evaluation | 2 | 1 | Mar 24, 2026 | |
| Scientific Review Feedback Generation | 2 | 1 | Mar 11, 2026 | |
| Reward Scoring | 2 | 2 | Apr 6, 2026 | |
| Prefill-stage hallucination risk detection | 2 | 1 | Mar 23, 2026 | |
| Prompt Hygiene Evaluation | 2 | 1 | Mar 23, 2026 | |
| Long-term preference alignment | 2 | 1 | Mar 27, 2026 | |
| High-level instruction execution | 2 | 1 | Mar 24, 2026 | |
| Multilingual Reward Modeling |
|---|
| 2 |
| 1 |
| Mar 12, 2026 |
| Multi-modal Dialogue | 2 | 1 | Mar 12, 2026 |
|---|
| Helpfulness Assessment | 2 | 2 | May 8, 2026 |
|---|
| Video Preference Evaluation | 2 | 2 | Apr 14, 2026 |
|---|
| Chat Performance | 2 | 1 | Mar 12, 2026 |
|---|
| Large Language Model Debiasing | 2 | 1 | Mar 20, 2026 |
|---|
| LLM Agent Task Completion | 2 | 2 | Jun 4, 2026 |
|---|
| Conflict Measurement | 2 | 1 | Mar 20, 2026 |
|---|
| Rationale Faithfulness Evaluation | 2 | 2 | May 12, 2026 |
|---|
| Multi-discipline Understanding | 2 | 4 | May 14, 2026 |
|---|
| Context Management | 2 | 2 | Jul 2, 2026 |
|---|
| Audio Instruction Following | 2 | 2 | May 20, 2026 |
|---|
| Shot-Language Understanding | 2 | 1 | Mar 20, 2026 |
|---|
| Web-based Reasoning | 2 | 2 | Jun 18, 2026 |
|---|
| Social Norm Dialogue Generation | 2 | 1 | Mar 13, 2026 |
|---|
| Zero-shot Reasoning and Knowledge | 2 | 2 | Apr 14, 2026 |
|---|
| Knowledge-focused evaluation | 2 | 1 | Feb 19, 2026 |
|---|