Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Federated Class-Incremental Learning | CIFAR-100 Quantity-based label imbalance | 42 | Apr 13, 2026 | ||
| Uncertainty Quantification | ImageNet 10 | 42 | Apr 13, 2026 | ||
| Uncertainty Quantification | CIFAR10 | 42 | Apr 13, 2026 | ||
| Uncertainty Quantification | FashionMNIST | 42 | Apr 13, 2026 | ||
| Fact Accuracy | Synthetic Phonebook Facts Power law Exponent beta = 1.0 |
| 42 |
| Apr 10, 2026 |
| Fact Accuracy | Synthetic Phonebook Facts Power law Exponent beta = 0 | 42 | Apr 10, 2026 |
|---|
| Visual Anomaly Detection | ViSA | 42 | Apr 9, 2026 |
|---|
| Model Fingerprinting Robustness | Structured Pruning Suspects Sheared-Llama | 42 | Apr 8, 2026 |
|---|
| Multiclass classification | TALENT | 42 | Jun 4, 2026 |
|---|
| AI-Generated Video Detection | VideoPhy 1.0 (test) | 42 | May 4, 2026 |
|---|
| AI-generated video detection | EvalCrafter | 42 | May 4, 2026 |
|---|
| AI-Generated Video Detection | VidProM | 42 | May 4, 2026 |
|---|
| Input Moderation | ToxicChat (test) | 42 | May 26, 2026 |
|---|
| Content Injection | Content Injection scenario | 42 | May 15, 2026 |
|---|
| Over Refusal | Over Refusal scenario | 42 | May 15, 2026 |
|---|
| Jailbreak | Jailbreak scenario | 42 | May 15, 2026 |
|---|
| Speed-of-Sound Reconstruction | In-silico C2 | 42 | Apr 3, 2026 |
|---|
| Medical Image Segmentation | JHU T0 (test) | 42 | Apr 2, 2026 |
|---|
| Medical Image Segmentation | MATAR T0 (test) | 42 | Apr 2, 2026 |
|---|
| Medical Image Segmentation | CHSF T0 (test) | 42 | Apr 2, 2026 |
|---|
| Search and Indexing Performance | Synthetic database of varying sizes (K) | 42 | Apr 2, 2026 |
|---|
| Anomaly Detection | SMD | 42 | Jun 23, 2026 |
|---|
| Text-to-Speech | Seed ZH | 42 | Jun 25, 2026 |
|---|
| Risk Scenario Evaluation | Google Workspace risk scenarios | 42 | Apr 1, 2026 |
|---|
| Zero-shot Language Understanding | ARC-Easy, ARC-Challenge, HellaSwag, LAMBADA, PIQA lm-eval 0.4.11 (test) | 42 | Mar 31, 2026 |
|---|