Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Visual-only Speech Recognition | LRS2 (test) | 77 | Jun 2, 2026 | ||
| Text Summarization | CNN/Daily Mail (test) | 77 | May 13, 2026 | ||
| Disparity Estimation | KITTI 2015 (test) | 77 | Feb 26, 2026 | ||
| Sequential Image Classification | pMNIST (test) |
| 77 |
| Feb 26, 2026 |
| Machine Comprehension | CNN (test) | 77 | Feb 26, 2026 |
|---|
| Image Classification | Office-10 + Caltech-10 | 77 | Feb 26, 2026 |
|---|
| Sentiment Classification | IMDB | 77 | Jun 26, 2026 |
|---|
| Graph Classification | NCI109 (test) | 77 | Jul 7, 2026 |
|---|
| Text-to-image generation | DPG-Bench | 77 | May 8, 2026 |
|---|
| Brain Tumor Segmentation | BraTS 2024 | 77 | Feb 26, 2026 |
|---|
| Image Quality Assessment | SPAQ (test) | 77 | Feb 26, 2026 |
|---|
| Medical Image Segmentation | Synapse | 77 | May 15, 2026 |
|---|
| Automatic Speech Recognition | LibriReplay-DOA (test) | 76 | Jun 30, 2026 |
|---|
| Jailbreak Attack | JailbreakBench | 76 | Jul 2, 2026 |
|---|
| Context-Aware Question Answering | TriState-Bench Sagr | 76 | Jun 11, 2026 |
|---|
| Context-Aware Question Answering | TriState-Bench Sres | 76 | Jun 11, 2026 |
|---|
| Context-Aware Question Answering | TriState-Bench Scor | 76 | Jun 11, 2026 |
|---|
| Error prediction | NCQA (test) | 76 | Jun 2, 2026 |
|---|
| Classification | Land Cover | 76 | Jun 2, 2026 |
|---|
| Long-term multivariate forecasting | Traffic (test) | 76 | Jun 26, 2026 |
|---|
| Visual Question Answering | TextVQA | 76 | Jul 3, 2026 |
|---|
| Long-context retrieval and aggregation | RULER 32k | 76 | May 27, 2026 |
|---|
| Long-context retrieval and aggregation | RULER 16k | 76 | May 27, 2026 |
|---|
| Long-context retrieval and aggregation | RULER 8k | 76 | May 27, 2026 |
|---|
| Long-context retrieval and aggregation | RULER 4k | 76 | May 27, 2026 |
|---|