Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Long Video Understanding | MLVU (dev) | 63 | Jun 11, 2026 | ||
| Node classification | CiteSeer | 63 | Jun 19, 2026 | ||
| Video Generation | Physics-IQ | 63 | Jun 5, 2026 | ||
| Image Classification | ImageYellow | 63 | May 21, 2026 | ||
| Image Classification | ImageMeow | 63 | May 21, 2026 | ||
| Pansharpening | QB (QuickBird) full-resolution (test) |
| 63 |
| Jun 17, 2026 |
| Image Classification | CIFAR-10 (Forget) | 63 | Apr 10, 2026 |
|---|
| Audio Reconstruction | AudioSet (eval) | 63 | Mar 24, 2026 |
|---|
| Question Answering | ASQA | 63 | Jun 30, 2026 |
|---|
| Open-Vocabulary Semantic Segmentation | ADE-847 | 63 | Apr 15, 2026 |
|---|
| Class-Incremental Learning | CIFAR-100 B0_Inc5 | 63 | May 29, 2026 |
|---|
| Harmful question-answering | BeaverTails HarmfulQA (1k and 10k samples) | 63 | Feb 26, 2026 |
|---|
| Concept Accuracy | Broden simplified (test) | 63 | Feb 26, 2026 |
|---|
| Mathematical Reasoning | CMATH | 63 | May 21, 2026 |
|---|
| Visual Speech Recognition | LRS3 | 63 | Mar 24, 2026 |
|---|
| Text Classification | Amazon | 63 | Mar 4, 2026 |
|---|
| Image-to-Text Retrieval | COCO 5K (test) | 63 | Jun 26, 2026 |
|---|
| Multimodal Understanding (Chinese) | MMBench Chinese | 63 | May 19, 2026 |
|---|
| Object Tracking | VisEvent (test) | 63 | Feb 26, 2026 |
|---|
| Conformal Prediction | ImageNet | 63 | Jun 8, 2026 |
|---|
| Traffic Forecasting | PEMS04 | 63 | Jun 1, 2026 |
|---|
| Backdoor Defense | CIFAR10 (train) | 63 | Feb 26, 2026 |
|---|
| Object Classification | ModelNet40 1k points | 63 | Feb 26, 2026 |
|---|
| Semantic Segmentation | Waymo Open Dataset (val) | 63 | Feb 26, 2026 |
|---|
| Language Modeling | WikiText-2 raw (test) | 63 | Jun 2, 2026 |
|---|