Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Machine Unlearning | CIFAR-10 1.0 (test) | 35 | May 19, 2026 | ||
| Mathematical reasoning | ASDiv Out of Distribution | 35 | Feb 26, 2026 | ||
| 3D Object Classification | ScanObjectNN | 35 | May 19, 2026 | ||
| Text Classification | SST-2 | 35 | Feb 26, 2026 | ||
| Conversational Question Answering | CoQA |
| 35 |
| Jun 9, 2026 |
| Semantic Segmentation | Deepglobe | 35 | Feb 26, 2026 |
|---|
| Multi-needle retrieval | NIAH (M) | 35 | Feb 26, 2026 |
|---|
| Mathematical Reasoning | SAT | 35 | Jun 9, 2026 |
|---|
| Node Classification | Squirrel filtered | 35 | Feb 26, 2026 |
|---|
| Code Generation | MBPP | 35 | Feb 26, 2026 |
|---|
| Grasp Generation | GRAB out-of-domain (test) | 35 | Feb 26, 2026 |
|---|
| Instruction Following | MT-Bench (test) | 35 | May 19, 2026 |
|---|
| Reasoning Segmentation | ReasonSeg | 35 | Jul 8, 2026 |
|---|
| Factor Analysis | LOOPR (in-sample) | 35 | Feb 26, 2026 |
|---|
| Action Recognition | NTU-RGB+D 120 (96/24) | 35 | May 13, 2026 |
|---|
| Text-to-Image Retrieval | MS-COCO 5-fold 1K (test) | 35 | Jul 3, 2026 |
|---|
| Image-to-Text Retrieval | MS-COCO 5-fold 1K (test) | 35 | Jul 3, 2026 |
|---|
| Aspect Sentiment Quadruple Prediction | FSQP (test) | 35 | Feb 26, 2026 |
|---|
| Jailbreak Defense | AdvBench PAIR attack | 35 | Feb 26, 2026 |
|---|
| Class-Incremental Learning | ImageNet-100 10T | 35 | Mar 13, 2026 |
|---|
| Commonsense Reasoning | BoolQ (test) | 35 | Jul 3, 2026 |
|---|
| Image Super-Resolution | DIV2K v1 (val) | 35 | Feb 26, 2026 |
|---|
| Question Answering | StrategyQA | 35 | Feb 26, 2026 |
|---|
| Open-QA Evaluation | EVOUNA-NaturalQuestions | 35 | Feb 26, 2026 |
|---|
| Chinese Spelling Correction | CSCD-NS | 35 | Feb 26, 2026 |
|---|