Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Video Perception | Perception (test) | 57 | Mar 9, 2026 | ||
| Agent Task | WebShop | 57 | Jul 7, 2026 | ||
| Question Answering | QA Zero-shot Average | 57 | Feb 26, 2026 | ||
| Mathematical Reasoning | MATH500 | 57 | Feb 26, 2026 | ||
| Reasoning | Checkmate-in-One | 57 | Mar 4, 2026 | ||
| Image Classification |
| COVIDx |
| 57 |
| Feb 27, 2026 |
| Object Detection | DOTA v1.5 (test) | 57 | Mar 18, 2026 |
|---|
| Speech Reconstruction | LibriTTS (test-other) | 57 | Jun 4, 2026 |
|---|
| Sequential Recommendation | Sports (test) | 57 | Jun 4, 2026 |
|---|
| Commonsense Reasoning | Commonsense Reasoning | 57 | Jun 2, 2026 |
|---|
| Class-Incremental Learning | Split ImageNet-R | 57 | May 13, 2026 |
|---|
| Recommendation | Amazon Sports (test) | 57 | Jun 30, 2026 |
|---|
| Recommendation | Amazon Baby (test) | 57 | Mar 16, 2026 |
|---|
| Speculative Decoding | Spec-Bench | 57 | May 28, 2026 |
|---|
| Open-Vocabulary Semantic Segmentation | PASCAL Context 59 (val) | 57 | Jul 7, 2026 |
|---|
| Counterfactual outcome prediction | MIMIC III semi-synthetic (800/200/200) | 57 | Feb 26, 2026 |
|---|
| Surface Reconstruction | Tanks and Temples | 57 | Apr 10, 2026 |
|---|
| Sequential Recommendation | Beauty (test) | 57 | Jun 4, 2026 |
|---|
| Text-to-audio retrieval | AudioCaps | 57 | Apr 28, 2026 |
|---|
| Domain Generalization | ImageNet variants (V2, S, A, R) (test) | 57 | May 27, 2026 |
|---|
| Image Classification | Tiny ImageNet | 57 | Feb 27, 2026 |
|---|
| Image Super-Resolution | B100 x3 (test) | 57 | May 19, 2026 |
|---|
| Causal Discovery | Synthetic Data | 57 | May 18, 2026 |
|---|
| Mathematical Reasoning | GSM8K | 57 | Feb 27, 2026 |
|---|
| Logical Reasoning | ProntoQA (test) | 57 | May 26, 2026 |
|---|