Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Depth Super-Resolution | ScanNet | 35 | Feb 26, 2026 | ||
| Multimodal Perception | HR-8K | 35 | Jun 30, 2026 | ||
| Multimodal Understanding | MMBench v1.1 (dev) | 35 | Apr 21, 2026 | ||
| Model Extraction Attack | CIFAR10 | 35 | Feb 26, 2026 | ||
| Safety Evaluation | XSTest | 35 | Jun 9, 2026 | ||
| Retrieval | Economic |
| 35 |
| Feb 26, 2026 |
| Evaluation-based Bias Reduction | Bias Reduction Benchmark (Evaluation) | 35 | Feb 26, 2026 |
|---|
| Memory-based Bias Reduction | Bias Reduction Benchmark Memory | 35 | Feb 26, 2026 |
|---|
| Needle-in-a-haystack | Needle-in-a-haystack 4x original context | 35 | Apr 14, 2026 |
|---|
| Code Reasoning | HumanE | 35 | Feb 26, 2026 |
|---|
| Math & Science | MATH 4-shot | 35 | Mar 4, 2026 |
|---|
| Knowledge Tracing | ASSIST09 (test) | 35 | May 12, 2026 |
|---|
| Classification | MultiRC | 35 | Jun 5, 2026 |
|---|
| Self-Harm Detection | JiraiBench 1.0 (test) | 35 | Feb 26, 2026 |
|---|
| Eating Disorder Detection | JiraiBench 1.0 (test) | 35 | Feb 26, 2026 |
|---|
| Overdose Detection | JiraiBench 1.0 (test) | 35 | Feb 26, 2026 |
|---|
| Summarization | CNN/DM | 35 | Feb 26, 2026 |
|---|
| Long Context Reasoning | AA-LCR | 35 | Jun 16, 2026 |
|---|
| Visual Question Answering | Encyclopedic-VQA Full | 35 | Feb 26, 2026 |
|---|
| Pluralistic Alignment | VITAL Overton | 35 | Feb 26, 2026 |
|---|
| Machine Translation | Wikinews-25 it->en | 35 | Feb 26, 2026 |
|---|
| Machine Translation | Wikinews-25 en->it | 35 | Feb 26, 2026 |
|---|
| Translation | FLORES-200 it-en (devtest) | 35 | Feb 26, 2026 |
|---|
| Translation | FLORES-200 en-it (devtest) | 35 | Feb 26, 2026 |
|---|
| Machine Translation | NTREX it->en 128 (test) | 35 | Feb 26, 2026 |
|---|