Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Code Reasoning | MBPP | 39 | Jun 18, 2026 | ||
| Reward Modeling | VLRewardBench (test) | 39 | May 12, 2026 | ||
| Machine Unlearning | RWKU Llama 3.1 8B (Forget Set) | 39 | Feb 26, 2026 | ||
| Tumor Segmentation | BraTS 23 | 39 | Apr 2, 2026 | ||
| Multiple Choice Question Answering | MedQA | 39 | Mar 16, 2026 | ||
| Image Classification |
| FEMNIST (LEAF) (test) |
| 39 |
| Feb 26, 2026 |
| Targeted Adversarial Attack | ImageNet | 39 | Feb 26, 2026 |
|---|
| Digital Watermarking | Blender and LLFF (test) | 39 | May 29, 2026 |
|---|
| Out-of-distribution detection | OpenOOD Far-OoD average v1.5 | 39 | Feb 26, 2026 |
|---|
| Out-of-distribution detection | OpenOOD Near-OoD average v1.5 | 39 | Feb 26, 2026 |
|---|
| Video Question Answering | PerceptionTest | 39 | Jul 7, 2026 |
|---|
| Image Super-resolution | Set5 2x | 39 | Jun 19, 2026 |
|---|
| Classification | PCAM | 39 | Feb 27, 2026 |
|---|
| Mathematical Reasoning | NuminaMath | 39 | Apr 20, 2026 |
|---|
| Multiple-Choice Question Answering | TruthfulQA MC1 | 39 | Apr 3, 2026 |
|---|
| Planning | nuScenes v1.0-trainval (val) | 39 | May 12, 2026 |
|---|
| Pixel-level Anomaly Detection | ColonDB | 39 | Apr 14, 2026 |
|---|
| Uncertainty Estimation | NaturalQA | 39 | May 1, 2026 |
|---|
| Automatic Speech Recognition | Earnings 22 | 39 | Jun 11, 2026 |
|---|
| Node Classification | Wisconsin H=0.19 | 39 | Feb 26, 2026 |
|---|
| Node Classification | Texas H=0.09 | 39 | Feb 26, 2026 |
|---|
| Node Classification | Cornell H=0.30 | 39 | Feb 26, 2026 |
|---|
| Node Classification | Citeseer H=0.74 | 39 | Feb 26, 2026 |
|---|
| Node Classification | Cora H=0.81 | 39 | Feb 26, 2026 |
|---|
| Graph Anomaly Detection | Elliptic | 39 | Apr 20, 2026 |
|---|