Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Multiple Choice Questions | Four-domain MCQ (test) | 43 | Jun 9, 2026 | ||
| Reasoning | ARC Challenge | 43 | Jun 9, 2026 | ||
| Face Verification | LFW | 43 | Jun 4, 2026 | ||
| Semantic Segmentation | ADE20K | 43 | Jul 7, 2026 | ||
| Time-Series Classification | WISDM 2 | 43 | Jun 23, 2026 | ||
| Question Answering |
| SearchQA |
| 43 |
| Jun 19, 2026 |
| Mathematical Reasoning | AIME 24 (test) | 43 | Jun 23, 2026 |
|---|
| Unconditional generation | LM1B sequence length 128 | 43 | May 11, 2026 |
|---|
| Language Modeling | WikiText-103 | 43 | May 19, 2026 |
|---|
| Synthetic Text Generation | MIMIC IV (test) | 43 | May 6, 2026 |
|---|
| Mathematical Reasoning | MATH | 43 | Jun 4, 2026 |
|---|
| Discriminative Object Hallucination | POPE MSCOCO Adversarial | 43 | Jun 1, 2026 |
|---|
| Semantic Segmentation | Pascal Context 60 with background | 43 | Apr 29, 2026 |
|---|
| Tabular Synthetic Data Generation | Default | 43 | May 19, 2026 |
|---|
| Safety Classification | ToxicChat (test) | 43 | May 26, 2026 |
|---|
| Proactive dialogue | ESConv | 43 | Jun 15, 2026 |
|---|
| Multi-modal Evaluation | MME | 43 | Jun 26, 2026 |
|---|
| Instance Segmentation | COCO | 43 | May 19, 2026 |
|---|
| Code Authorship Attribution | CoDET-M4 | 43 | Apr 21, 2026 |
|---|
| Code Authorship Attribution | LLMAuthorBench | 43 | Apr 21, 2026 |
|---|
| Mathematical Reasoning | Math-500 | 43 | Apr 21, 2026 |
|---|
| Question Answering | TruthfulQA | 43 | May 15, 2026 |
|---|
| Mathematical reasoning | GSM8K (test) | 43 | Jun 29, 2026 |
|---|
| Hallucination Detection | POPE | 43 | Jun 26, 2026 |
|---|
| Mathematical Reasoning | MathVista mini | 43 | Jul 7, 2026 |
|---|