Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Underwater Image Enhancement | UIEB (test) | 38 | May 15, 2026 | ||
| Document Classification | Reuters | 38 | Mar 4, 2026 | ||
| Skeleton-based Action Recognition | NTU RGB+D Cross-View (CV) 1.0 | 38 | Feb 26, 2026 | ||
| Multi-hop Question Answering | HotpotQA fullwiki setting (dev) |
| 38 |
| Mar 4, 2026 |
| Weakly-supervised object localization | CUB-200 2011 (test) | 38 | Feb 26, 2026 |
|---|
| 3D Face Reconstruction | NoW face challenge (test) | 38 | Feb 26, 2026 |
|---|
| Fine-grained object category discovery | Stanford Cars (test) | 38 | Feb 26, 2026 |
|---|
| Text-to-Image Synthesis | COCO (test) | 38 | Feb 26, 2026 |
|---|
| Machine Reading Comprehension | RACE | 38 | Feb 26, 2026 |
|---|
| 3D Shape Retrieval | ModelNet40 (test) | 38 | Feb 26, 2026 |
|---|
| Summarization | Gigaword (test) | 38 | Feb 26, 2026 |
|---|
| Summarization | Gigaword | 38 | Feb 27, 2026 |
|---|
| Digit Classification | USPS -> MNIST | 38 | Feb 27, 2026 |
|---|
| Landmark Prediction | MAFL (test) | 38 | Feb 26, 2026 |
|---|
| Part-of-Speech Tagging | TWEEBANK V2 (test) | 38 | Feb 26, 2026 |
|---|
| Pedestrian Detection | CrowdHuman | 38 | Feb 26, 2026 |
|---|
| Facial Landmark Detection | AFLW Front | 38 | Feb 26, 2026 |
|---|
| Age Estimation | FG-NET | 38 | Jun 9, 2026 |
|---|
| Age Estimation | MORPH S2 (Setting II) | 38 | Feb 26, 2026 |
|---|
| Video Classification | Charades | 38 | Feb 27, 2026 |
|---|
| Action Recognition | ActivityNet (test) | 38 | Feb 26, 2026 |
|---|
| Image Classification | Places 205-way (test) | 38 | Feb 26, 2026 |
|---|
| Image Quality Assessment | CSIQ (full) | 38 | Feb 26, 2026 |
|---|
| Short Text Clustering | StackOverflow | 38 | Feb 26, 2026 |
|---|
| Short Text Clustering | SearchSnippets | 38 | Feb 26, 2026 |
|---|