Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Event Prediction | TAXI | 47 | Jun 9, 2026 | ||
| Pose Estimation | HiRoom | 47 | Jul 7, 2026 | ||
| Text Continuation | WikiText-103 512-token continuation (test) | 47 | May 11, 2026 | ||
| Common Sense Reasoning | HellaSwag | 47 | Jun 2, 2026 | ||
| Interactive Decision Making | ALFWorld Seen | 47 | Jul 7, 2026 | ||
| Classification | SPAM (test) |
| 47 |
| Apr 7, 2026 |
| Agent Safety | AuraGen | 47 | Apr 7, 2026 |
|---|
| Multivariate Time Series Classification | UEA multivariate time-series archive (test) | 47 | May 28, 2026 |
|---|
| Image Restoration | MiO100 AgenticIR setting (Group B) | 47 | May 22, 2026 |
|---|
| Image Restoration | MiO100 AgenticIR setting (Group A) | 47 | May 22, 2026 |
|---|
| Safety Evaluation | FigStep | 47 | May 13, 2026 |
|---|
| Model Fingerprinting | MNIST | 47 | Mar 27, 2026 |
|---|
| Model Fingerprinting | CIFAR-10 | 47 | Mar 27, 2026 |
|---|
| Mathematical Reasoning | Math Benchmarks Average | 47 | May 27, 2026 |
|---|
| Image Classification | PCAM | 47 | Apr 23, 2026 |
|---|
| Visual Perception | AI2D | 47 | May 25, 2026 |
|---|
| Long-context retrieval and synthetic reasoning | RULER | 47 | Mar 24, 2026 |
|---|
| Cross-view Geo-localization (Satellite to Drone) | SUES-200 300m altitude | 47 | May 12, 2026 |
|---|
| Emotional Mimicry Intensity Estimation | Hume-Vidmimic2 (val) | 47 | Mar 17, 2026 |
|---|
| Inductive dynamic link prediction | Can. Parl. (Inductive) | 47 | Jun 4, 2026 |
|---|
| Classification | car | 47 | May 12, 2026 |
|---|
| Chart Understanding | ChartQA | 47 | Apr 28, 2026 |
|---|
| LLM Pretraining | C4 | 47 | May 22, 2026 |
|---|
| Coding | Eval+ | 47 | Jun 16, 2026 |
|---|
| Speculative Decoding | SpecBench | 47 | May 27, 2026 |
|---|