Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Graph Classification | PROTEINS | 72 | Jun 1, 2026 | ||
| Retrieval Question Answering | RQA | 72 | May 22, 2026 | ||
| Prior Estimation | FASHION | 72 | May 22, 2026 | ||
| Prior Estimation | CIFAR | 72 | May 22, 2026 | ||
| Prior Estimation | MNIST | 72 | May 22, 2026 | ||
| Membership Inference Attack | WikiMIA K = 16 |
| 72 |
| May 22, 2026 |
| Image Classification | ImageNette ImageNet-1K (test) | 72 | May 21, 2026 |
|---|
| Targeted Attack on Image Captioning | Frontier MLLM Evaluation Set | 72 | May 20, 2026 |
|---|
| Job Shop Scheduling | TA | 72 | May 19, 2026 |
|---|
| Language Modeling | WikiText-2 | 72 | May 19, 2026 |
|---|
| Visual Active Search | DOTA 10x10 settings (test) | 72 | May 18, 2026 |
|---|
| Visual Active Search | xView single-target category (test) | 72 | May 18, 2026 |
|---|
| Prompt Optimization | XSum | 72 | May 15, 2026 |
|---|
| Best feasible prompt identification | CNN/DailyMail (test) | 72 | May 15, 2026 |
|---|
| Off-policy evaluation | OBP library avg over 50 datasets | 72 | May 14, 2026 |
|---|
| CN vs. MCI vs. AD classification | ADNI | 72 | May 14, 2026 |
|---|
| Offline Reinforcement Learning under Gravity Shift | MuJoCo Walker2d | 72 | May 14, 2026 |
|---|
| Ultrasound Image Denoising | In vivo Ultrasound 8 sub-apertures | 72 | May 14, 2026 |
|---|
| Ultrasound Image Denoising | In vivo Ultrasound 8 sub-apertures | 72 | May 14, 2026 |
|---|
| Runtime Performance Evaluation | BEEBS (Entire Suite) | 72 | May 12, 2026 |
|---|
| Unsafe Robustness | AdvBench | 72 | May 12, 2026 |
|---|
| Unsafe Robustness | HarmBench | 72 | May 12, 2026 |
|---|
| Jailbreak Robustness | HarmBench | 72 | May 12, 2026 |
|---|
| Unsafe Robustness | JailbreakBench | 72 | May 12, 2026 |
|---|
| Model Editing | ZsRE | 72 | May 12, 2026 |
|---|