Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Multi-task Language Understanding | MMLU Biz | 41 | Jun 4, 2026 | ||
| Single-turn Jailbreak Attack | HarmBench single-turn | 41 | Jun 4, 2026 | ||
| Model Inversion Defense | CelebA 64x64 | 41 | Jun 2, 2026 | ||
| Clustering | CIFAR-10 | 41 | Jun 2, 2026 | ||
| Hateful Meme Detection | FBHM (test) | 41 | Jun 1, 2026 | ||
| AMC23 |
| 41 |
| Jun 16, 2026 |
| Pairwise evaluation | BIGGEN | 41 | Jun 1, 2026 |
|---|
| Language Modeling | LAMBADA | 41 | May 29, 2026 |
|---|
| Response Harmfulness Detection | SafeRLHF | 41 | Jun 1, 2026 |
|---|
| Object Detection | OD-LVIS | 41 | May 27, 2026 |
|---|
| Image Generation | ImageNet-1K 256x256 (train val) | 41 | May 27, 2026 |
|---|
| Code Generation | HumanEval | 41 | Jul 7, 2026 |
|---|
| Physics Reasoning | Public Physics Benchmarks (GPQA, SciBench, PhysReason) (test) | 41 | Jul 7, 2026 |
|---|
| Task-invariant subspace estimation | Simulated data (d=100, s=5) | 41 | May 14, 2026 |
|---|
| Mathematical Reasoning | AIME 26 | 41 | May 21, 2026 |
|---|
| Web Navigation Question Answering | WebWalkerQA | 41 | Jun 11, 2026 |
|---|
| Object Probing | POPE Adversarial | 41 | Jun 29, 2026 |
|---|
| Object Probing | POPE Random | 41 | Jun 29, 2026 |
|---|
| Multi-task Evaluation | PostTrainBench | 41 | Jun 2, 2026 |
|---|
| Mathematical Reasoning | Omni-Math | 41 | Jun 23, 2026 |
|---|
| Image Classification | EuroSAT | 41 | Jun 23, 2026 |
|---|
| Language Modeling | The Pile (val) | 41 | Jun 24, 2026 |
|---|
| TreeSHAP Explanation | Synthetic d=10 | 41 | May 8, 2026 |
|---|
| 3D Object Detection | nuScenes LiDAR Beamsreduce | 41 | Jun 1, 2026 |
|---|
| Anomaly Detection | ADBench Tabular (aggregated across 47 datasets) | 41 | Jun 30, 2026 |
|---|