Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Question Answering | RQA | 130 | May 25, 2026 | ||
| Class-conditional generation | ImageNet 256 x 256 1k (val) | 130 | Jul 3, 2026 | ||
| Open-loop planning | nuScenes | 130 | Jun 30, 2026 | ||
| Text-to-Image Generation | GenEval 1.0 (test) |
| 130 |
| May 26, 2026 |
| Image Classification | 8 vision benchmarks (roSAT, GTSRB, MNIST, RESISC, Aircraft, SVHN, etc.) | 130 | Feb 26, 2026 |
|---|
| Comprehensive Multi-modal Evaluation | MME | 130 | Jul 3, 2026 |
|---|
| Language Modeling | Proof-pile | 130 | Jun 15, 2026 |
|---|
| Language Understanding | MMLU-Pro | 130 | Jun 30, 2026 |
|---|
| Image Generation | ImageNet 64x64 | 130 | Jun 12, 2026 |
|---|
| Image Captioning | COCO | 130 | Mar 30, 2026 |
|---|
| Information Retrieval | BEIR (test) | 130 | Jun 16, 2026 |
|---|
| Video Action Recognition | HMDB51 | 130 | May 13, 2026 |
|---|
| Image Classification | PACS | 130 | May 29, 2026 |
|---|
| Image Captioning | NoCaps | 130 | Jun 24, 2026 |
|---|
| Language Modeling | Penn Treebank (PTB) (test) | 130 | Apr 23, 2026 |
|---|
| Node Classification | Cora standard (test) | 130 | Feb 26, 2026 |
|---|
| Deepfake detection | DFDC (test) | 130 | Apr 24, 2026 |
|---|
| Video Summarization | SumMe | 130 | Feb 26, 2026 |
|---|
| Action Recognition | UCF101 (Split 1) | 130 | Jun 23, 2026 |
|---|
| Information Visual Question Answering | InfoVQA (test) | 130 | Mar 27, 2026 |
|---|
| Question Answering | OpenBookQA (OBQA) (test) | 130 | Feb 26, 2026 |
|---|
| Video Compression (Y Component) | JVET Sequences Classes A1, A2, B, C, D, E, F, G, M (test) | 129 | Jun 2, 2026 |
|---|
| Image Classification | SUN397 | 129 | Jun 16, 2026 |
|---|
| Multi-view Clustering | LandUse-21 | 129 | Jun 4, 2026 |
|---|
| Spatial Reasoning | Viewspatial | 129 | Jun 11, 2026 |
|---|