Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Semantic Segmentation | ScanNet v2 (test) | 248 | Feb 27, 2026 | ||
| Image Deblurring | HIDE (test) | 248 | Jun 23, 2026 | ||
| Object Classification | ScanObjectNN OBJ_BG | 248 | May 19, 2026 | ||
| Image Generation |
| ImageNet 256x256 (train) |
| 247 |
| Jun 23, 2026 |
| Instance Segmentation | Cityscapes (val) | 247 | Jun 11, 2026 |
|---|
| Multi-hop Question Answering | 2WikiMultiHopQA (test) | 247 | Jun 30, 2026 |
|---|
| Question Answering | ARC Easy | 246 | Jul 7, 2026 |
|---|
| Mathematical Reasoning | GSM8K | 246 | Mar 20, 2026 |
|---|
| Image Classification | CIFAR-10 | 246 | Mar 10, 2026 |
|---|
| Long-Tailed Image Classification | ImageNet-LT (test) | 246 | May 12, 2026 |
|---|
| Multivariate long-term forecasting | Electricity | 245 | Jun 9, 2026 |
|---|
| Action Recognition | Kinetics 400 (test) | 245 | Mar 20, 2026 |
|---|
| Stereo Matching | KITTI 2015 (test) | 245 | Jun 30, 2026 |
|---|
| Commonsense Reasoning | Commonsense Reasoning (BoolQ, PIQA, SIQA, HellaS., WinoG., ARC-e, ARC-c, OBQA) | 245 | Jul 2, 2026 |
|---|
| Knowledge Graph Question Answering | CWQ | 245 | Jul 2, 2026 |
|---|
| Mathematical Reasoning | MGSM | 245 | Jun 26, 2026 |
|---|
| Jailbreak Attack | SafeBench | 245 | May 20, 2026 |
|---|
| Red-Teaming | HarmBench | 244 | Apr 23, 2026 |
|---|
| Language Modeling | PG-19 | 244 | Jun 15, 2026 |
|---|
| Referring Video Object Segmentation | Ref-YouTube-VOS (val) | 244 | Mar 31, 2026 |
|---|
| Text-to-Image Retrieval | MS-COCO 5K (test) | 244 | Apr 3, 2026 |
|---|
| Referring Expression Comprehension | RefCOCO+ (testB) | 244 | Jul 7, 2026 |
|---|
| Scene Text Recognition | IIIT5K (test) | 244 | Feb 26, 2026 |
|---|
| Image Classification | Oxford Flowers 102 | 244 | Jul 8, 2026 |
|---|
| Mathematical Reasoning | AIME 2024 | 243 | Jun 12, 2026 |
|---|