Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Feature Visualization Evaluation | ImageNet | 37 | Feb 26, 2026 | ||
| Mathematical Reasoning | AMC 23 | 37 | Mar 4, 2026 | ||
| Time Series Forecasting | GIFT-Eval | 37 | Jul 3, 2026 | ||
| multi-coil MRI reconstruction | fastMRI Brain 8x acceleration multi-coil | 37 | May 20, 2026 | ||
| multi-coil MRI reconstruction | fastMRI Brain multi-coil 4x acceleration |
| 37 |
| May 20, 2026 |
| Voice Cloning | SEED-TTS-Eval ZH (test) | 37 | Jun 8, 2026 |
|---|
| Task-Oriented Dialogue | MultiWOZ 2.0 (test) | 37 | Feb 26, 2026 |
|---|
| Math Problem Solving | AIME 24 | 37 | Jun 23, 2026 |
|---|
| Machine Unlearning | CIFAR-10 Random Forget 10% (train) | 37 | May 13, 2026 |
|---|
| Long-context understanding | RULER 64K | 37 | May 19, 2026 |
|---|
| Music Genre Classification | FMA small (test) | 37 | Mar 17, 2026 |
|---|
| Causal Discovery | Tübingen | 37 | Apr 8, 2026 |
|---|
| Image Classification | DomainBed | 37 | Jun 18, 2026 |
|---|
| Mathematical Reasoning | AIME25 | 37 | Apr 10, 2026 |
|---|
| Assortment optimization | SUSHI Preference Data | 37 | Feb 26, 2026 |
|---|
| Question Answering | ARC-E 0-shot | 37 | Apr 23, 2026 |
|---|
| Mathematical Visual Question Answering | MathVerse | 37 | Apr 10, 2026 |
|---|
| Code Generation | HumanEval v1 (test) | 37 | May 27, 2026 |
|---|
| Sequential Recommendation | Yelp (test) | 37 | Apr 20, 2026 |
|---|
| Multimodal Recommendation | Amazon Clothing (test) | 37 | May 4, 2026 |
|---|
| Generative Recommendation | Yelp | 37 | May 13, 2026 |
|---|
| Medical Image Classification | Kvasir | 37 | May 22, 2026 |
|---|
| Spatial Reasoning | CV-Bench-3D | 37 | May 20, 2026 |
|---|
| Visual Search | HR-Bench 4K | 37 | Jun 30, 2026 |
|---|
| Multi-contrast MRI Reconstruction | M4raw | 37 | Jun 16, 2026 |
|---|