Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Partial Multi-Label Learning | reference | 36 | Apr 13, 2026 | ||
| Partial Multi-Label Learning | arts | 36 | Apr 13, 2026 | ||
| Text-based Person Re-identification | ICFG-PEDES | 36 | Jun 2, 2026 | ||
| Federated Continual Learning | CIFAR10 3 Tasks | 36 | Apr 13, 2026 | ||
| Super-Resolution | DIV8K | 36 | Apr 13, 2026 | ||
| Simultaneous Speech Translation |
| MuST-C en-de v1.0 (test) |
| 36 |
| Apr 13, 2026 |
| Simultaneous Speech Translation | MuST-C en-ru v1.0 (test) | 36 | Apr 13, 2026 |
|---|
| Simultaneous Speech Translation | MuST-C en-ro 1.0 (test) | 36 | Apr 13, 2026 |
|---|
| Simultaneous Speech Translation | MuST-C en-pt v1.0 (test) | 36 | Apr 13, 2026 |
|---|
| Simultaneous Speech Translation | MuST-C en-nl 1.0 (test) | 36 | Apr 13, 2026 |
|---|
| Question Answering | Dolly Closed QA | 36 | Apr 10, 2026 |
|---|
| Question Answering | SQuAD v2 | 36 | Apr 10, 2026 |
|---|
| Image Compression | LSUN (test) | 36 | Apr 10, 2026 |
|---|
| Image Compression | CelebA-HQ (test) | 36 | Apr 10, 2026 |
|---|
| Mathematical Reasoning | MATH 500 | 36 | Apr 10, 2026 |
|---|
| Long-context Question Answering | En.QA | 36 | Apr 10, 2026 |
|---|
| Long-context Question Answering | NarrativeQA | 36 | Apr 10, 2026 |
|---|
| Long-context Question Answering | MFQA En | 36 | Apr 10, 2026 |
|---|
| Long-context Question Answering | 2WikiMQA | 36 | Apr 10, 2026 |
|---|
| Regression | Insurance | 36 | Apr 10, 2026 |
|---|
| Classification | Magic | 36 | Apr 10, 2026 |
|---|
| Image Classification | CIFAR-10 (test) | 36 | Apr 10, 2026 |
|---|
| Image Classification | CIFAR-100 | 36 | Apr 10, 2026 |
|---|
| Downstream Task Accuracy via Paradigm Routing | Downstream Tasks (test) | 36 | Apr 10, 2026 |
|---|
| Multi-hop Question Answering | MuSiQue | 36 | May 13, 2026 |
|---|