Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Mathematical Reasoning | MathVision | 89 | Jul 7, 2026 | ||
| Medical Image Segmentation | FLARE 2021 | 89 | May 19, 2026 | ||
| Commonsense Question Answering | WinoGrande | 89 | Jun 30, 2026 | ||
| Lifelong Model Editing | ZsRE | 89 | May 13, 2026 | ||
| Math Reasoning | SVAMP |
| 89 |
| Jun 9, 2026 |
| Mathematical Reasoning | GSM8K 8-shot | 89 | May 11, 2026 |
|---|
| Science Question Answering | OpenBookQA | 89 | Jun 24, 2026 |
|---|
| Code Generation | MBPP | 89 | May 20, 2026 |
|---|
| Tabular Classification | 75 Tabular Classification Datasets (test) | 89 | Feb 26, 2026 |
|---|
| Video Reasoning | Video-Holmes | 89 | Jul 2, 2026 |
|---|
| Image-level Anomaly Detection | VisA (test) | 89 | Jun 15, 2026 |
|---|
| Text-based Visual Question Answering | VQAText | 89 | Jul 3, 2026 |
|---|
| Spatial Reasoning | CV-Bench | 89 | Jun 16, 2026 |
|---|
| LLM alignment evaluation | AlpacaEval 2 | 89 | May 28, 2026 |
|---|
| Deep Research Report Generation | DeepResearch Bench | 89 | Jun 9, 2026 |
|---|
| Mathematical Reasoning | BRUMO25 | 89 | Jun 16, 2026 |
|---|
| Multimodal Reasoning | MathVista | 89 | Jun 26, 2026 |
|---|
| Reward Modeling | RewardBench v1.0 (test) | 89 | Mar 4, 2026 |
|---|
| Multi-Task Learning | PASCAL Context | 89 | Jun 4, 2026 |
|---|
| Audio-Visual Video Parsing | LLP (test) | 89 | May 13, 2026 |
|---|
| Image Classification | ImageNet-Sketch | 89 | May 29, 2026 |
|---|
| Code Generation | LiveCodeBench | 89 | Feb 26, 2026 |
|---|
| Visual Question Answering | ScienceQA image | 89 | Jul 3, 2026 |
|---|
| Reading Comprehension | C3 | 89 | May 26, 2026 |
|---|
| Base-to-New Generalization | StanfordCars | 89 | Jul 2, 2026 |
|---|