Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Image Super-Resolution | Set14 | 289 | Feb 26, 2026 | ||
| Scene Text Recognition | SVT (test) | 289 | Feb 26, 2026 | ||
| Mathematical Reasoning | AIME | 288 | Mar 5, 2026 | ||
| Video Question Answering | ActivityNet-QA (test) |
| 288 |
| Mar 20, 2026 |
| Common-sense Reasoning | COPA | 288 | Jun 29, 2026 |
|---|
| Video Anomaly Detection | UCF-Crime | 288 | Jul 2, 2026 |
|---|
| Panoptic Segmentation | Cityscapes (val) | 288 | May 19, 2026 |
|---|
| Reasoning Segmentation | ReasonSeg (test) | 287 | Jul 8, 2026 |
|---|
| Instruction Following | MT-Bench | 287 | Jun 29, 2026 |
|---|
| Language Generation | WikiText2 | 287 | May 19, 2026 |
|---|
| Image Classification | CIFAR-100 (test) | 287 | Mar 6, 2026 |
|---|
| Image Classification | ImageNet-ReaL | 287 | Jul 9, 2026 |
|---|
| Dynamic multi-objective optimization | DF and FDA benchmark suites (DF1-DF14, FDA1-FDA5) MIGD values Modified | 285 | Apr 2, 2026 |
|---|
| Node Classification | Photo | 285 | Jun 30, 2026 |
|---|
| Time Series Forecasting | Electricity | 285 | Jun 17, 2026 |
|---|
| Image Classification | CIFAR10 (test) | 284 | Feb 27, 2026 |
|---|
| Referring Expression Segmentation | RefCOCO+ (val) | 284 | Jun 4, 2026 |
|---|
| Reward Modeling | RewardBench | 284 | Jul 7, 2026 |
|---|
| Jailbreak Attack | AdvBench | 283 | Jun 5, 2026 |
|---|
| Visual Question Answering | OKVQA | 283 | Feb 27, 2026 |
|---|
| Image Classification | Waterbirds | 283 | Jul 2, 2026 |
|---|
| Question Answering | SciQ | 283 | Apr 8, 2026 |
|---|
| Question Answering | ARC-C | 283 | Jun 8, 2026 |
|---|
| Referring Image Segmentation | RefCOCO (val) | 283 | Jun 23, 2026 |
|---|
| Commonsense Reasoning | Commonsense Reasoning (BoolQ, PIQA, SIQA, HellaS., WinoG., ARC-e, ARC-c, OBQA) (test) | 283 | Jul 1, 2026 |
|---|