Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Depth Super-Resolution | TOFDSR | 40 | May 12, 2026 | ||
| Depth Super-Resolution | RGB-D-D | 40 | May 12, 2026 | ||
| GUI Automation | OSWorld Verified (test) | 40 | May 27, 2026 | ||
| Point Cloud Classification | ScanObjectNN v1 (test) | 40 | Feb 26, 2026 | ||
| Jailbreak Attack Resistance Evaluation | Jailbreak | 40 | Feb 26, 2026 |
| Quantum Key Distribution Runtime Analysis | Simulated QKD Environment | 40 | Feb 26, 2026 |
|---|
| Adversarial Attack | LVLM Evaluation Set | 40 | Feb 26, 2026 |
|---|
| Mathematical Reasoning | Minerva-Math | 40 | Mar 4, 2026 |
|---|
| Open-Domain Question Answering | TriviaQA | 40 | Feb 26, 2026 |
|---|
| Open-Domain Question Answering | NaturalQuestions (NQ) | 40 | Feb 26, 2026 |
|---|
| Document Question Answering | MMLongBench-Doc | 40 | Jun 30, 2026 |
|---|
| Factuality Correction | ASKHIST | 40 | Feb 26, 2026 |
|---|
| Few-shot Learning | SAMSum | 40 | Feb 27, 2026 |
|---|
| Question Answering | NarrativeQA | 40 | Feb 27, 2026 |
|---|
| Alignment | Weather | 40 | Feb 26, 2026 |
|---|
| Clinical Prediction | MIMIC-IV | 40 | Feb 26, 2026 |
|---|
| Common Phrase Labeling | CPL | 40 | Feb 26, 2026 |
|---|
| Entity-aware Sentence Alignment | ESA-MT | 40 | Feb 26, 2026 |
|---|
| Grammar Error Correction | GEC | 40 | Feb 26, 2026 |
|---|
| Named Entity Recognition | NER | 40 | Feb 26, 2026 |
|---|
| CPL | CPL | 40 | Feb 26, 2026 |
|---|
| ESA-MT | ESA-MT | 40 | Feb 26, 2026 |
|---|
| Grammatical Error Correction | GEC | 40 | Feb 26, 2026 |
|---|
| Speech Recognition | CommonVoice | 40 | Feb 26, 2026 |
|---|
| Reasoning | AIME 25 | 40 | Feb 26, 2026 |
|---|