Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Depth Super-Resolution / Completion | NYU v2 (test) | 36 | Feb 26, 2026 | ||
| Motion Deblurring | ImageNet | 36 | May 14, 2026 | ||
| Real-world Super-Resolution | RealSR | 36 | May 12, 2026 | ||
| Visual Question Answering | FloodNet | 36 | Feb 26, 2026 | ||
| Visual Question Answering | GQA | 36 | Feb 26, 2026 | ||
| VQA v2 |
| 36 |
| Feb 26, 2026 |
| Style Unlearning | UnlearnCanvas | 36 | Jun 1, 2026 |
|---|
| Panoptic Segmentation | COCO 2017 | 36 | Feb 26, 2026 |
|---|
| Watermarking Robustness | LVLM Evaluation Set (test) | 36 | Feb 26, 2026 |
|---|
| Smoke Segmentation | MSSDataset Medium | 36 | Feb 26, 2026 |
|---|
| Adversarial Attack Imperceptibility | ImageNet | 36 | Mar 25, 2026 |
|---|
| Image Steganalysis | BOSSBase 1.01 (test) | 36 | Jul 8, 2026 |
|---|
| Multiclass Classification | NSL-KDD | 36 | Apr 1, 2026 |
|---|
| Runtime Attestation Overhead Evaluation | CoreMark Pro | 36 | May 12, 2026 |
|---|
| Membership Inference Attack | CNN-DM | 36 | Feb 26, 2026 |
|---|
| Membership Inference Attack | MedInstruct | 36 | Feb 26, 2026 |
|---|
| Membership Inference Attack | HealthcareMagic | 36 | Feb 26, 2026 |
|---|
| Intrusion Detection | CAR HACKING dataset | 36 | Feb 26, 2026 |
|---|
| IoT Intrusion Detection | N-BaIoT Scenario 2 | 36 | Feb 26, 2026 |
|---|
| Vulnerability Reasoning | Vulnerability Reasoning CWE-Guided Prompt (test) | 36 | Feb 26, 2026 |
|---|
| Vulnerability Reasoning | Vulnerability Reasoning Basic Prompt (test) | 36 | Feb 26, 2026 |
|---|
| Subject inference attack | zsRE batch-edit tasks | 36 | Feb 26, 2026 |
|---|
| Subject inference attack | CounterFact batch-edit tasks | 36 | Feb 26, 2026 |
|---|
| Safety alignment evaluation | Llama-Guard | 36 | Feb 26, 2026 |
|---|
| Quality Estimation | PAWS-X (test) | 36 | Feb 26, 2026 |
|---|