Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Object Detection | Pascal-5i 2010 (Novel Split 1) | 54 | Feb 26, 2026 | ||
| Visual Search | V* benchmark | 54 | May 28, 2026 | ||
| Synthetic Image Generation | Camelyon | 54 | May 29, 2026 | ||
| Synthetic Image Generation | CelebA | 54 | May 29, 2026 | ||
| Imprinting Attack | COCO | 54 | Feb 26, 2026 | ||
| Image Classification |
| N-MNIST binary |
| 54 |
| Feb 26, 2026 |
| Mathematical Reasoning | AIME 2024 | 54 | Apr 21, 2026 |
|---|
| Multi-Hop Question Answering | HotpotQA | 54 | Apr 21, 2026 |
|---|
| Mathematical Reasoning | Minerva Math | 54 | Apr 9, 2026 |
|---|
| Sentiment Analysis | CR | 54 | Feb 26, 2026 |
|---|
| Instruction Following and Reasoning | Low-resource languages evaluation suite (am, arz, ars, as, ast, az, ba, bn, bo, ceb, cv, cy, fo, ga, gd, gl, gn, ha, ht, ig, jv, kmr, sdh, ky, lb, lo, lus, mg, mi, mn, mt, ny, oc, pap, ps, rn, rw, sd, si, sm, sn, st, su, sw, te, tg, ti, tk, tt, ug, xh, yi, yo, zu) | 54 | Feb 26, 2026 |
|---|
| Prompt classification | Aegis | 54 | Jul 8, 2026 |
|---|
| Mathematical Reasoning | HMMT Feb 2025 | 54 | Jun 4, 2026 |
|---|
| Scientific Reasoning | GPQA diamond | 54 | May 12, 2026 |
|---|
| Machine-Generated Text Detection | TruthfulQA | 54 | May 18, 2026 |
|---|
| Mathematical Reasoning | Math500 Level 5 | 54 | Feb 26, 2026 |
|---|
| General Question Answering | TriviaQA | 54 | Apr 10, 2026 |
|---|
| Time Series Imputation | Exchange | 54 | Feb 26, 2026 |
|---|
| Mathematical Reasoning | MetaMathQA | 54 | Feb 26, 2026 |
|---|
| Mathematical Reasoning | AIME 2024 | 54 | Feb 26, 2026 |
|---|
| Structural Consistency | Consistency Evaluation Dataset (N=720) (test) | 54 | Feb 26, 2026 |
|---|
| Consistency Analysis | Consistency Analysis 80 cases (test) | 54 | Feb 26, 2026 |
|---|
| Semantic Consistency | Consistency evaluation suite N=720 (test) | 54 | Feb 26, 2026 |
|---|
| Long-context language modeling | LongBench (test) | 54 | Jun 30, 2026 |
|---|
| Operations Research Problem Solving | StructuredOR (test) | 54 | Feb 26, 2026 |
|---|