Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Video-to-Text Retrieval | MSVD | 125 | Jun 18, 2026 | ||
| Image Classification | ImageNet (val) | 125 | Mar 9, 2026 | ||
| Image Generation | ImageNet 256x256 (test) | 125 | Jun 26, 2026 | ||
| Multi-label Classification |
|---|
| PASCAL VOC 2007 (test) |
| 125 |
| Feb 26, 2026 |
| Semantic Segmentation | PASCAL-Context 59 class (val) | 125 | Feb 27, 2026 |
|---|
| Language Modeling | One Billion Word Benchmark (test) | 125 | Jun 1, 2026 |
|---|
| Multimodal Understanding | SEED-2-Plus | 125 | Jun 9, 2026 |
|---|
| Skeleton-based Action Recognition | NTU 60 (X-view) | 125 | Jun 11, 2026 |
|---|
| Survival Prediction | TCGA-STAD | 125 | Jun 23, 2026 |
|---|
| Novel View Synthesis | NeRF Synthetic | 125 | Jul 7, 2026 |
|---|
| Mathematical Reasoning | Minerva Math | 124 | Jun 23, 2026 |
|---|
| Multimodal Mathematical Reasoning | MathVista MINI | 124 | Jun 9, 2026 |
|---|
| Mathematical reasoning | MATH500 | 124 | Jun 11, 2026 |
|---|
| Text Classification | BoolQ | 124 | Jun 5, 2026 |
|---|
| Video Question Answering | EgoSchema subset | 124 | May 12, 2026 |
|---|
| Question Answering | NarrativeQA | 124 | Jun 2, 2026 |
|---|
| Automatic Speech Recognition | LibriSpeech clean | 124 | Jun 18, 2026 |
|---|
| Visual Question Answering | TextVQA (test) | 124 | Feb 27, 2026 |
|---|
| Image Denoising | SIDD | 124 | Jun 23, 2026 |
|---|
| Image Retrieval | Revisited Oxford (ROxf) (Medium) | 124 | Feb 26, 2026 |
|---|
| Image Reconstruction | ImageNet1K (val) | 124 | Jun 4, 2026 |
|---|
| Image Classification | CIFAR-10 | 124 | Feb 26, 2026 |
|---|
| Fine-grained Image Classification | Stanford Dogs (test) | 124 | Apr 8, 2026 |
|---|
| Multi-Label Classification | NUS-WIDE (test) | 124 | May 28, 2026 |
|---|
| Question Answering | CommonsenseQA (CSQA) | 124 | Feb 26, 2026 |
|---|