Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Skeleton-based Action Recognition | NTU 60 (X-sub) | 227 | Jun 11, 2026 | ||
| Image Classification | EuroSAT | 226 | May 1, 2026 | ||
| Multimodal Understanding | SEED | 226 | Jun 19, 2026 | ||
| Multivariate long-term series forecasting | Traffic (test) |
| 226 |
| May 28, 2026 |
| Graph Regression | ZINC (test) | 226 | Jun 12, 2026 |
|---|
| 3D Object Detection | KITTI Car (test) | 226 | Mar 19, 2026 |
|---|
| Depth Estimation | NYU Depth v2 | 226 | Jun 12, 2026 |
|---|
| Visual Question Answering | VQAv2 | 226 | Jul 1, 2026 |
|---|
| Image Classification | CIFAR-10 (test) | 225 | Jun 26, 2026 |
|---|
| Visual Question Answering | SimpleVQA | 225 | Jun 26, 2026 |
|---|
| Commonsense Reasoning | ARC-C | 225 | Jun 30, 2026 |
|---|
| Image Classification | CIFAR-100 LT | 225 | Jun 4, 2026 |
|---|
| Action Recognition | UCF-101 | 225 | May 25, 2026 |
|---|
| Image Classification | CIFAR-10 NAS-Bench-201 (test) | 225 | May 22, 2026 |
|---|
| Action Recognition | HMDB51 | 225 | Feb 26, 2026 |
|---|
| Interactive Segmentation | GrabCut | 225 | Feb 26, 2026 |
|---|
| Open-loop planning | nuScenes (val) | 225 | May 26, 2026 |
|---|
| Instruction Following | InstructBench | 224 | Mar 26, 2026 |
|---|
| Code Generation | HumanEval | 224 | Jun 4, 2026 |
|---|
| Mathematical Reasoning | AMC | 224 | Jun 9, 2026 |
|---|
| Clustering | CIFAR-10 (test) | 224 | Jun 24, 2026 |
|---|
| Surface Normal Estimation | NYU v2 (test) | 224 | Jul 1, 2026 |
|---|
| Video Question-Answering | LongVideoBench | 224 | Jul 7, 2026 |
|---|
| Graduate-level Question Answering | GPQA | 224 | Jun 19, 2026 |
|---|
| Robot Manipulation | LIBERO | 223 | Jun 30, 2026 |
|---|