Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Clinical quality indicator checking | CMQCIC-Bench v1 (test) | 121 | Feb 26, 2026 | ||
| Hierarchical Image Classification | CUB-200-2011 | 120 | Jul 8, 2026 | ||
| Hierarchical Image Classification | Aircraft (test) | 120 | Jul 8, 2026 | ||
| Shortest Path | 5x5 Grid Graph (test) | 120 | Jun 19, 2026 | ||
| 3D Path Planning | AUV LAR 1.0 (Evaluation Sets) | 120 |
| Jun 5, 2026 |
| 3D Path Planning | AUV LAR (test) | 120 | Jun 5, 2026 |
|---|
| Numerical Optimization | CEC 2014 | 120 | Jun 5, 2026 |
|---|
| Membership Inference | VL-MIA DALL·E | 120 | May 14, 2026 |
|---|
| Empirical coverage estimation | RiverSwim | 120 | May 13, 2026 |
|---|
| Query Routing | LongBench OOD v2 | 120 | May 12, 2026 |
|---|
| Hyperspectral Image Classification | WHU-Hi-LongKou (WHLK) | 120 | Apr 30, 2026 |
|---|
| Targeted Adversarial Attack | ImageNet (val) | 120 | Apr 23, 2026 |
|---|
| Targeted Adversarial Attack | ImageNet ILSVRC2012 (val) | 120 | Apr 23, 2026 |
|---|
| Covariance matrix estimation | Block covariance structure | 120 | Mar 31, 2026 |
|---|
| Covariance Estimation | Factor covariance structure synthetic | 120 | Mar 31, 2026 |
|---|
| Traveling Salesman Problem | TSPLib Large-scale instances | 120 | Mar 31, 2026 |
|---|
| Online Bin Packing | Online BPP (test) | 120 | Mar 25, 2026 |
|---|
| Reward Modeling | RMB | 120 | Apr 10, 2026 |
|---|
| Uncertainty Quantification | Average of 6 datasets | 120 | Feb 26, 2026 |
|---|
| Multi-Robot Path Planning | TSPLIB | 120 | Feb 26, 2026 |
|---|
| Click-Through Rate Prediction | Industrial | 120 | May 26, 2026 |
|---|
| Instance Segmentation | PanNuke 19 tissue types (three-fold cross-validation) | 120 | Feb 26, 2026 |
|---|
| Temporal Grounding | Charades-STA | 120 | Jul 7, 2026 |
|---|
| Long Video Understanding | Video-MME Long | 120 | Jul 8, 2026 |
|---|
| Long-context language modeling evaluation | FDA (test) | 120 | Feb 26, 2026 |
|---|