Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Short-form generation | Short-form generation ID | 38 | Apr 14, 2026 | ||
| Text Classification | AGNews | 38 | May 21, 2026 | ||
| Question Answering | WMDP Biology | 38 | Apr 14, 2026 | ||
| Question Answering | WMDP Cyber QA | 38 | Apr 14, 2026 | ||
| Recommendation | Yelp | 38 | May 26, 2026 | ||
| Dialogue Response Generation | MSC |
| 38 |
| Apr 10, 2026 |
| Dialogue Response Generation | Chronicle | 38 | Apr 10, 2026 |
|---|
| Safety Evaluation | Alpaca | 38 | Apr 21, 2026 |
|---|
| Sarcasm Detection | MMSD 1.0 (test) | 38 | Apr 9, 2026 |
|---|
| Homomorphic Encryption Kernel Latency | HE Kernels N=2^16 | 38 | Apr 7, 2026 |
|---|
| Anomaly Reasoning | MMAD | 38 | Jul 7, 2026 |
|---|
| Remote Sensing Referring Expression Segmentation (RSRES) | RISBench (test) | 38 | Jul 7, 2026 |
|---|
| Jailbreak Defense | Manual (IJP) | 38 | Apr 3, 2026 |
|---|
| Time Series Classification | 18 UEA datasets Regular | 38 | Apr 3, 2026 |
|---|
| Classification | BasicMotions 50% Missing | 38 | Jul 2, 2026 |
|---|
| Multimodal Understanding | LVLM Evaluation Suite (AI2D, DocVQA, InfoVQA, MMBench, MME, MMMU, SciQA, TextVQA, MMStar, POPE) | 38 | Apr 2, 2026 |
|---|
| Object Hallucination Mitigation on Generative Tasks | AMBER | 38 | Apr 21, 2026 |
|---|
| Math Reasoning | MATH 500 | 38 | May 18, 2026 |
|---|
| Adversarial Purification | CIFAR-100 | 38 | Mar 31, 2026 |
|---|
| Remote Object Grounding | REVERIE (test unseen) | 38 | May 27, 2026 |
|---|
| Remote Object Grounding | REVERIE (val unseen) | 38 | May 27, 2026 |
|---|
| General AI Assistant tasks | GAIA | 38 | May 19, 2026 |
|---|
| Text-to-Video Retrieval | MSRVTT | 38 | Mar 27, 2026 |
|---|
| Object Detection | KAIST (test) | 38 | Mar 27, 2026 |
|---|
| Object Detection | FLIR-ADAS (test) | 38 | Mar 27, 2026 |
|---|