Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Node Classification | Cora | 65 | Feb 26, 2026 | ||
| Semantic Segmentation | WeedMap (test) | 65 | Mar 10, 2026 | ||
| Camera Pose Estimation | TUM | 65 | Jul 2, 2026 | ||
| Action Recognition | SSv2 Small | 65 | May 28, 2026 | ||
| Object-goal navigation | HM3D-OVON Unseen (val) |
| 65 |
| Jun 9, 2026 |
| Multi-image Understanding | MMIU | 65 | Apr 14, 2026 |
|---|
| Zero-shot Classification | CIFAR100 | 65 | May 20, 2026 |
|---|
| Spatial Reasoning | MMSI-Bench | 65 | May 29, 2026 |
|---|
| Machine Translation Evaluation | WMT Metrics Shared Task 2024 | 65 | Mar 16, 2026 |
|---|
| Classification | WSC | 65 | Jun 5, 2026 |
|---|
| Language Modeling | LM1B | 65 | Jun 30, 2026 |
|---|
| Long-context understanding | RULER | 65 | Apr 16, 2026 |
|---|
| Mathematical Reasoning | AMC (test) | 65 | May 28, 2026 |
|---|
| Harmful Request Defense | AdvBench | 65 | Jun 25, 2026 |
|---|
| Multi-Hop Question Answering | Multi-Hop QA (HotpotQA, 2Wiki, Musique, Bamboogle) (test) | 65 | Jun 11, 2026 |
|---|
| Mathematical Reasoning | AIME 25 | 65 | Feb 26, 2026 |
|---|
| Deepfake Detection | FaceForensics++ (test) | 65 | May 11, 2026 |
|---|
| Multimodal Reasoning | MMBench | 65 | Jul 1, 2026 |
|---|
| Medical Visual Question Answering | VQA-RAD (test) | 65 | Jun 12, 2026 |
|---|
| Image Classification | CIFAR-100 Dir-0.1 | 65 | Apr 21, 2026 |
|---|
| Medical Image Classification | MedMnist BloodMnist (test) | 65 | Feb 26, 2026 |
|---|
| Topic Classification | AGNEWS | 65 | Jun 11, 2026 |
|---|
| Video Question Answering | ActivityNet-QA zero-shot (test) | 65 | Jun 23, 2026 |
|---|
| Video Question Answering | MSRVTT-QA zero-shot (test) | 65 | Jun 23, 2026 |
|---|
| Function-level Code Generation | HumanEval+ augmented (test) | 65 | May 11, 2026 |
|---|