Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Sequence Classification | IMDB | 64 | Feb 26, 2026 | ||
| Sequence Classification | Yahoo | 64 | Feb 26, 2026 | ||
| Sequence Classification | Huffpost low-resource (test) | 64 | Feb 26, 2026 | ||
| Model Inversion Defense | CelebA | 64 | Feb 27, 2026 | ||
| Continual Learning | ImageNet-R 10 tasks | 64 | Jul 7, 2026 | ||
| CC |
| 64 |
| Apr 3, 2026 |
| Watermark Detection | Llama-2-7b-chat-hf 10 samples UMD watermarking (test) | 64 | Feb 26, 2026 |
|---|
| No-Reference Image Quality Assessment | LIVEC | 64 | Feb 26, 2026 |
|---|
| Face Anti-Spoofing | CASIA-CeFA (C), PADISI (P), CASIA-SURF (S), and WMCA (W) Protocol 2, Missing D | 64 | Jul 7, 2026 |
|---|
| Language Modeling | Pre-training (val) | 64 | Jun 16, 2026 |
|---|
| Video Object Segmentation | MOSE (val) | 64 | Jul 7, 2026 |
|---|
| Semantic Segmentation | A-847 | 64 | Apr 28, 2026 |
|---|
| Univariate Time-series Forecasting | ETTm2 (test) | 64 | Feb 26, 2026 |
|---|
| Regression | synthetic datasets #2 (test) | 64 | Feb 26, 2026 |
|---|
| Regression | synthetic datasets #1 1.0 (test) | 64 | Feb 26, 2026 |
|---|
| Few-Shot Semantic Segmentation | FSS-1000 | 64 | Feb 26, 2026 |
|---|
| Semantic Segmentation | VOC | 64 | Jun 15, 2026 |
|---|
| Language Modeling | arXiv | 64 | Jun 30, 2026 |
|---|
| OOD Detection | OpenImage-O | 64 | Jun 17, 2026 |
|---|
| Video Grounding | QVHighlights (test) | 64 | Feb 26, 2026 |
|---|
| Depth Super-Resolution | Middlebury (test) | 64 | May 12, 2026 |
|---|
| Vision-and-Language Navigation | REVERIE seen (val) | 64 | May 25, 2026 |
|---|
| Vectorized Map Construction | nuScenes v1.0 (val) | 64 | Feb 26, 2026 |
|---|
| Multilingual Mathematical Reasoning | MGSM | 64 | Jun 30, 2026 |
|---|
| Traveling Salesman Problem (TSP) | TSP n=100 10K instances (test) | 64 | Jun 30, 2026 |
|---|