Loading the SOTA2 catalog…
SOTA2 Research · tasks
Explore the problems researchers are working on and the benchmarks used to measure progress.
| Task Name | Domain | Benchmarks | Papers | Last Update |
|---|---|---|---|---|
| Fairness and Utility Evaluation | 1 | 1 | Apr 10, 2026 | |
| Story Editing | 1 | 1 | Feb 19, 2026 | |
| Mathematics Evaluation | 1 | 1 | Apr 10, 2026 | |
| Logic Task | 1 | 1 | Apr 10, 2026 | |
| Storybook Generation Evaluation | 1 | 1 | Feb 19, 2026 | |
| Medical LLM Alignment | 1 | 1 | Apr 10, 2026 | |
| Multimodal Large Language Model Inference Efficiency | 1 | 1 | Feb 19, 2026 | |
| Dialogue voice design | 1 | 1 | Apr 10, 2026 | |
| Fact Memorization |
| 1 |
| 1 |
| Apr 10, 2026 |
| Language Model Accuracy Evaluation | 1 | 1 | Apr 13, 2026 |
|---|
| Dialogue-based Mathematical Problem Solving | 1 | 1 | Feb 18, 2026 |
|---|
| Representation Injection Performance | 1 | 1 | Apr 13, 2026 |
|---|
| Myopic choice evaluation | 1 | 1 | Apr 13, 2026 |
|---|
| Task Representation Accuracy | 1 | 1 | Apr 13, 2026 |
|---|
| Output Length Maximization | 1 | 1 | Feb 18, 2026 |
|---|
| Aggregated Downstream Evaluation | 1 | 1 | Apr 13, 2026 |
|---|
| Prompt-embedding Optimization | 1 | 2 | Apr 14, 2026 |
|---|
| Language-conditioned long-horizon robotic manipulation | 1 | 3 | Jul 7, 2026 |
|---|
| Chinese Subjective Alignment | 1 | 1 | Feb 18, 2026 |
|---|
| Reasoning and comprehension | 1 | 1 | Apr 13, 2026 |
|---|
| Large Multimodal Model Inference Efficiency | 1 | 1 | Feb 19, 2026 |
|---|
| Dialogue Faithfulness Evaluation | 1 | 1 | Apr 13, 2026 |
|---|
| Tool-augmented Graph Reasoning | 1 | 1 | Apr 13, 2026 |
|---|
| Text-to-Image Instruction Following | 1 | 2 | Apr 15, 2026 |
|---|
| Video Instruction Editing | 1 | 1 | Apr 13, 2026 |
|---|