Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Scene Text Recognition | SVT 647 (test) | 101 | Feb 26, 2026 | ||
| Super-Resolution | BSDS100 (test) | 101 | Apr 24, 2026 | ||
| Facial Landmark Detection | AFLW Full | 101 | Feb 26, 2026 | ||
| Text Detection | ICDAR MLT 2017 (test) | 101 | Feb 26, 2026 | ||
| Question Answering | NewsQA (dev) |
| 101 |
| Feb 26, 2026 |
| Offline Reinforcement Learning | D4RL walker2d-random | 101 | May 11, 2026 |
|---|
| Semantic Segmentation | COCO Object (val) | 101 | May 7, 2026 |
|---|
| Image Quality Assessment | KADID-10k (test) | 101 | Apr 23, 2026 |
|---|
| Referring Video Segmentation | MeViS | 101 | May 15, 2026 |
|---|
| Visual Question Answering | OKVQA (val) | 101 | Feb 26, 2026 |
|---|
| Ktrans Synthesis | Clinical Brain-Tumor Cohort (test) | 100 | Jun 25, 2026 |
|---|
| Speech Watermarking Robustness | AISHELL-3 (test) | 100 | Jun 16, 2026 |
|---|
| Load Forecasting | Midea dataset | 100 | Jun 9, 2026 |
|---|
| Clustering | CUB V=2, K=10 (test) | 100 | Jun 4, 2026 |
|---|
| Syntactic Correctness | C++ | 100 | Jun 2, 2026 |
|---|
| Copyright Verification | CIFAR-10 (test) | 100 | May 14, 2026 |
|---|
| Calibration | HMDA (test) | 100 | May 12, 2026 |
|---|
| Object Hallucination Evaluation | POPE (popular) | 100 | Jul 3, 2026 |
|---|
| Code Generation | HumanEval 0-shot | 100 | Jun 26, 2026 |
|---|
| Image-to-Brain Retrieval | Natural Scenes Dataset (NSD) | 100 | Apr 15, 2026 |
|---|
| Class-conditioned image generation | ImageNet-1K 1.0 (test val) | 100 | Mar 5, 2026 |
|---|
| Agentic Reasoning | τ-Bench | 100 | Feb 27, 2026 |
|---|
| Common Sense Reasoning | PIQA | 100 | May 7, 2026 |
|---|
| Binary Code Similarity Detection | 10 victim functions and 10,366 dissimilar functions (test) | 100 | Feb 26, 2026 |
|---|
| Automatic Speech Recognition | OpenASR | 100 | Jul 9, 2026 |
|---|