Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Time Series Forecasting | ETTh2 | 66 | Mar 4, 2026 | ||
| Membership Inference | WikiMIA 32 tokens 1.0 | 66 | Feb 26, 2026 | ||
| Class Unlearning | CIFAR-10 | 66 | Jun 24, 2026 | ||
| Tabular Anomaly Detection | Ionosphere | 66 | May 12, 2026 | ||
| Network Cost Optimization | Exp-3 Large-scale Network Instances | 66 | Feb 26, 2026 | ||
| Retrieval |
|---|
| MS MARCO v1 |
| 66 |
| May 1, 2026 |
| Multimodal Recommendation | Amazon Baby (test) | 66 | May 4, 2026 |
|---|
| Hyperspectral Image Super-Resolution | PaviaU (test) | 66 | Apr 14, 2026 |
|---|
| Image Captioning Evaluation | Nebula | 66 | Jun 30, 2026 |
|---|
| Perceptual Quality Assessment | HPE-Bench 1.0 (test) | 66 | Feb 26, 2026 |
|---|
| Visual Question Answering | VQA | 66 | May 19, 2026 |
|---|
| Drone-to-Satellite Retrieval | SUES-200 300m | 66 | Apr 3, 2026 |
|---|
| Drone-to-Satellite Retrieval | SUES-200 200m | 66 | Apr 3, 2026 |
|---|
| Multi-discipline Multimodal Understanding | MMMU Pro | 66 | May 18, 2026 |
|---|
| SAR Image Classification | MSTAR (test) | 66 | Apr 3, 2026 |
|---|
| Fingerprint Removal | LLM Fingerprinting Evaluation Alpaca-GPT4-52k | 66 | Feb 26, 2026 |
|---|
| Machine Unlearning | MNIST | 66 | Jun 24, 2026 |
|---|
| Timing Attack | N-MNIST integer | 66 | Feb 26, 2026 |
|---|
| Influence Estimation | Benchmarks Budgets k=1, 5, 10, 25 (Aggregated) | 66 | Feb 26, 2026 |
|---|
| Acoustic Consistency | SALMon | 66 | Apr 21, 2026 |
|---|
| Fake News Detection | FAKE NEWS | 66 | Feb 26, 2026 |
|---|
| Mathematical Reasoning | AIME’25, AIME’24, AMC’23, and MATH500 Average (test) | 66 | Feb 26, 2026 |
|---|
| Question Answering | 2WikiMQA | 66 | Jun 4, 2026 |
|---|
| Language Understanding | MMLU CF | 66 | Apr 28, 2026 |
|---|
| Machine Translation | En-Es document-level | 66 | Mar 27, 2026 |
|---|