Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Few-shot Image Classification | miniImageNet meta (test) | 46 | Feb 26, 2026 | ||
| Concealed Object Detection | NC4K | 46 | Feb 26, 2026 | ||
| Model Selection | DTD | 46 | Feb 26, 2026 | ||
| Recommendation | REDIAL (test) | 46 | Feb 26, 2026 | ||
| Language Modeling | BLLIP-LG (test) | 46 | May 18, 2026 | ||
| Sequential Recommendation | Amazon Toys and Games (test) |
| 46 |
| Jun 15, 2026 |
| Zero-shot Accuracy | ARC Challenge | 46 | Feb 26, 2026 |
|---|
| Optical Flow | MVSEC 1.0 (indoor_flying3) | 46 | Mar 5, 2026 |
|---|
| Optical Flow | MVSEC 1.0 (indoor_flying2) | 46 | Mar 5, 2026 |
|---|
| Order Completion | E-I environment mu=0.5 | 45 | Jul 8, 2026 |
|---|
| Order Completion | E-I environment mu=0.6 | 45 | Jul 8, 2026 |
|---|
| LoRA Routing | NQ-DomainLoRA PMD-BENCH | 45 | Jul 7, 2026 |
|---|
| Question Answering | PaperQA | 45 | Jul 7, 2026 |
|---|
| Dataset Distillation | CIFAR-100 | 45 | Jul 2, 2026 |
|---|
| Clustering | Lymphoma synthetic missingness | 45 | Jul 1, 2026 |
|---|
| Long-Context Reasoning | InfiniteBench | 45 | Jul 1, 2026 |
|---|
| Mathematical Reasoning | UMath | 45 | Jun 30, 2026 |
|---|
| Long-term multivariate forecasting | Electricity (test) | 45 | Jun 26, 2026 |
|---|
| Closed-set Classification | CIFAR-100 (50-50) | 45 | Jun 26, 2026 |
|---|
| Closed-set Classification | CIFAR-100 (20 / 80) | 45 | Jun 26, 2026 |
|---|
| Closed-set Classification | CIFAR-10 (6 4) | 45 | Jun 26, 2026 |
|---|
| Coding | CodeContest | 45 | Jun 18, 2026 |
|---|
| Coding | LiveCodeBench Aug 24 – Jan 25 | 45 | Jun 18, 2026 |
|---|
| Emotional Understanding | STEU | 45 | Jun 18, 2026 |
|---|
| Knowledge Distillation | CIFAR-100 (val) | 45 | Jun 18, 2026 |
|---|