Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Long-term Time-Series Forecasting | Sinewave (SIN) synthetic (test) | 36 | Jun 2, 2026 | ||
| Image Classification | ImageNet-1K | 36 | Jun 2, 2026 | ||
| Molecular system generation | Alanine-Dipeptide | 36 | Jun 2, 2026 | ||
| HRTF Upsampling | HRTF Dataset | 36 | Jun 2, 2026 | ||
| Equation Discovery | Gravitational potential | 36 | Jun 2, 2026 | ||
| Navier–Stokes |
| 36 |
| Jun 2, 2026 |
| Equation Discovery | Heat conduction | 36 | Jun 2, 2026 |
|---|
| Coding Reasoning | NSCLC | 36 | Jun 2, 2026 |
|---|
| Coding Reasoning | Frontal Cortex | 36 | Jun 2, 2026 |
|---|
| Mirage Detection | VQA Mirage Detection Chest, Pathology, Natural, Document, Infographic (test) | 36 | Jun 2, 2026 |
|---|
| Question Answering | ATE-Bench Q&A suite (Q1–Q12) | 36 | Jun 1, 2026 |
|---|
| Privacy and Utility Evaluation | PEEP | 36 | Jun 1, 2026 |
|---|
| Privacy and Utility Evaluation | PasswordEval | 36 | Jun 1, 2026 |
|---|
| Mathematical Reasoning | AIME '25 | 36 | Jun 1, 2026 |
|---|
| Mathematical Reasoning | MATH500 | 36 | Jun 1, 2026 |
|---|
| Corpus Composition Inference | LLMScan Mid-Grained 1.0 (test) | 36 | May 29, 2026 |
|---|
| Mathematical Reasoning | AIME25 | 36 | Jun 9, 2026 |
|---|
| Factual QA | NQ-Open | 36 | Jun 2, 2026 |
|---|
| Question Answering | ARC Challenge | 36 | May 28, 2026 |
|---|
| Question Answering | ARC Easy | 36 | May 28, 2026 |
|---|
| Safety Evaluation | Refusal | 36 | May 28, 2026 |
|---|
| Truthful Question Answering | TruthfulQA | 36 | May 28, 2026 |
|---|
| Out-of-sample Mean Squared Error Estimation | Spiked covariance model N = 200 | 36 | May 28, 2026 |
|---|
| Argument Quality Assessment | Webis-ArgQuality 20 | 36 | May 28, 2026 |
|---|
| Object Hallucination Assessment | COCO 500 images | 36 | May 28, 2026 |
|---|