Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Hallucination Detection | PopQA | 97 | Jun 2, 2026 | ||
| Zero-shot Reasoning | Evaluation Suite Zero-shot (OpenbookQA, ARC-e, ARC-c, WinoGrande, HellaSwag, PIQA, MathQA) | 97 | Jun 29, 2026 | ||
| General AI Assistant Task | GAIA (val) | 97 | May 29, 2026 | ||
| Table Question Answering | TABMWP | 97 | Apr 21, 2026 | ||
| Mathematical reasoning | GSM8K |
| 97 |
| Mar 4, 2026 |
| Long Video Understanding | VideoMME | 97 | Jul 2, 2026 |
|---|
| Semantic Segmentation | LoveDA | 97 | Apr 20, 2026 |
|---|
| Jailbreak Defense | PAIR | 97 | Apr 14, 2026 |
|---|
| Person Re-identification | LTCC cloth-changing | 97 | Jun 11, 2026 |
|---|
| Planning | nuScenes (val) | 97 | May 25, 2026 |
|---|
| Text-to-Speech | LibriSpeech clean (test) | 97 | Jun 18, 2026 |
|---|
| Image Classification | Food101 (test) | 97 | Jun 16, 2026 |
|---|
| Node classification | Questions (test) | 97 | Jun 4, 2026 |
|---|
| Image Classification | CIFAR-10 standard (test) | 97 | Feb 26, 2026 |
|---|
| Offline Reinforcement Learning | D4RL Medium-Replay HalfCheetah | 97 | May 11, 2026 |
|---|
| Image Super-Resolution | Manga109 x2 (test) | 97 | Jun 19, 2026 |
|---|
| Grayscale Image Denoising | Urban100 | 97 | Mar 30, 2026 |
|---|
| Video Classification | Kinetics-400 (test) | 97 | Feb 26, 2026 |
|---|
| Monocular Depth Estimation | NYU-Depth v2 (official) | 97 | Apr 30, 2026 |
|---|
| Image Denoising | SIDD Benchmark | 97 | May 26, 2026 |
|---|
| Super-Resolution | Urban100 x3 | 97 | May 19, 2026 |
|---|
| Image Inpainting | FFHQ (test) | 97 | Jul 7, 2026 |
|---|
| Action Recognition | Kinetics-600 | 97 | May 25, 2026 |
|---|
| Unsupervised Domain Adaptation | DomainNet (test) | 97 | Feb 26, 2026 |
|---|
| Relation Extraction | TACRED | 97 | Feb 26, 2026 |
|---|