Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Localization | ImageNet-1k (val) | 79 | Feb 26, 2026 | ||
| Super-Resolution | Set14 | 79 | Feb 26, 2026 | ||
| Visual Object Tracking | VastTrack | 78 | Jul 2, 2026 | ||
| Single Object Tracking | LaSOT (test) | 78 | Jul 2, 2026 | ||
| Generic Object Tracking | GOT-10K | 78 | Jul 2, 2026 | ||
| Language Modeling | DCLM (val) |
| 78 |
| Jul 7, 2026 |
| Joint VaR-ES forecasting | US equity (test) | 78 | Jun 4, 2026 |
|---|
| Visual Place Recognition | Oxford RobotCar (Dusk) | 78 | Jun 1, 2026 |
|---|
| 3D Adversarial Attack | ModelNet40 | 78 | May 20, 2026 |
|---|
| Image Classification | CIFAR10-DVS | 78 | Jul 1, 2026 |
|---|
| Interactive Decision Making | ScienceWorld | 78 | Jun 11, 2026 |
|---|
| Image Classification | Aircraft | 78 | Jul 8, 2026 |
|---|
| Safety Evaluation | XSTest Safe | 78 | May 7, 2026 |
|---|
| Code Generation | HumanEval | 78 | Jun 17, 2026 |
|---|
| Truthfulness | TruthfulQA | 78 | Jun 30, 2026 |
|---|
| In-hospital mortality prediction | eICU (test) | 78 | May 19, 2026 |
|---|
| Image Classification | ImageNet-1k | 78 | Jun 2, 2026 |
|---|
| Language Modeling | Lambada | 78 | Jun 30, 2026 |
|---|
| Forecasting | ETTh1 | 78 | Jun 9, 2026 |
|---|
| Classification | TOX_171 | 78 | May 6, 2026 |
|---|
| Joint Depth Super-Resolution and Denoising | NYU v2 (test) | 78 | Mar 24, 2026 |
|---|
| Mathematical Reasoning | Minerva | 78 | Jun 23, 2026 |
|---|
| Image Classification | CIFAR-100 Non-IID Dir(0.3) 2009 (test) | 78 | May 4, 2026 |
|---|
| Precipitation Nowcasting | WeatherBench (test) | 78 | Mar 17, 2026 |
|---|
| Mathematical Reasoning | AIME 24 | 78 | Jun 23, 2026 |
|---|