Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Sentiment Analysis | CMU-MOSEI (test) | 96 | May 20, 2026 | ||
| Classification | COIL-20 | 96 | May 28, 2026 | ||
| Emotion Detection | EmoryNLP (test) | 96 | Feb 26, 2026 | ||
| Domain Adaptation | OFFICE | 96 | Feb 26, 2026 | ||
| Image-to-Text Retrieval | MSCOCO (1K test) |
| 96 |
| Mar 31, 2026 |
| Reading Comprehension | DROP | 96 | May 21, 2026 |
|---|
| Face Anti-Spoofing | Idiap Replay-Attack OCM → I | 96 | Feb 26, 2026 |
|---|
| Fine-grained classification | Stanford Cars | 96 | Jun 25, 2026 |
|---|
| Polyp Segmentation | ColonDB | 96 | Jul 2, 2026 |
|---|
| Visual Question Answering | Vizwiz (val) | 96 | Jul 3, 2026 |
|---|
| Flowshop Scheduling | VRF hard-large n > 200 (test) | 95 | Jun 16, 2026 |
|---|
| Code Generation | MBPP | 95 | Jul 1, 2026 |
|---|
| Plan Cost Reduction | Planning Domains Individual Domains IPC | 95 | Apr 1, 2026 |
|---|
| Zero-shot Evaluation | Zero-shot Tasks Average | 95 | Jul 9, 2026 |
|---|
| Aspect-based Sentiment Analysis | ASAP (test) | 95 | Mar 31, 2026 |
|---|
| Image-to-image retrieval | Food101 | 95 | Jul 8, 2026 |
|---|
| Watermark Detection | C4 | 95 | May 29, 2026 |
|---|
| Question Answering | ARC-E and PIQA (test) | 95 | Feb 26, 2026 |
|---|
| Speaker Verification | VoxCeleb-E | 95 | Jul 9, 2026 |
|---|
| Math reasoning | AMC | 95 | Mar 31, 2026 |
|---|
| Anomaly Detection | MSL | 95 | Mar 20, 2026 |
|---|
| Video Understanding | LVBench | 95 | Jul 7, 2026 |
|---|
| Reasoning-based text-to-image generation | WISE | 95 | Jun 12, 2026 |
|---|
| Short-term Forecasting | PEMS08 | 95 | Jun 26, 2026 |
|---|
| Short-term Forecasting | PEMS04 | 95 | Jun 26, 2026 |
|---|