Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Software Engineering Issue Resolution | SWE-bench Verified | 110 | Jun 24, 2026 | ||
| Credit Card Fraud Detection | BankSim | 110 | May 13, 2026 | ||
| Jailbreaking | HarmBench | 110 | Jun 4, 2026 | ||
| Multimodal Understanding | MMMU | 110 | Jun 19, 2026 | ||
| Open-vocabulary semantic segmentation | ADE20K |
| 110 |
| Jun 16, 2026 |
| Image Classification | CIFAR-10 non-IID (test) | 110 | May 8, 2026 |
|---|
| Table Fact Verification | TabFact | 110 | Jun 8, 2026 |
|---|
| Semantic Segmentation | ScanNet Environments 1-10 Sequence Split | 110 | Feb 26, 2026 |
|---|
| Real-world Visual Understanding | RealWorldQA | 110 | May 11, 2026 |
|---|
| World Knowledge Image Generation | WISE | 110 | May 22, 2026 |
|---|
| Sign Language Translation | PHOENIX14T (test) | 110 | Jul 7, 2026 |
|---|
| Semantic Segmentation | Potsdam | 110 | May 28, 2026 |
|---|
| Class Incremental Learning | ImageNet-A | 110 | May 22, 2026 |
|---|
| Traveling Salesman Problem | TSP-500 (test) | 110 | Jun 18, 2026 |
|---|
| Super-Resolution | ImageNet (test) | 110 | Jun 11, 2026 |
|---|
| Multi-hop Question Answering | Bamboogle (test) | 110 | Jun 4, 2026 |
|---|
| Vision-and-Language Navigation | REVERIE unseen (test) | 110 | May 27, 2026 |
|---|
| Natural Language Inference | MNLI (matched) | 110 | Feb 26, 2026 |
|---|
| Image Classification | CIFAR-100 (test) | 110 | Mar 17, 2026 |
|---|
| Recipe-to-Image Retrieval | Recipe1M 1k setup (test) | 110 | Apr 20, 2026 |
|---|
| Graph Classification | MolHIV | 110 | Jul 3, 2026 |
|---|
| Image Quality Assessment | CSIQ (test) | 110 | Apr 23, 2026 |
|---|
| Referring Video Object Segmentation | JHMDB-Sentences (test) | 110 | Jun 26, 2026 |
|---|
| Image Classification | Flowers | 110 | Jun 2, 2026 |
|---|
| Text Classification | RTE | 110 | Jun 5, 2026 |
|---|