Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Visual Reasoning | GQA | 96 | Jun 24, 2026 | ||
| MAP inference | UAI Inference Competition 2022 | 96 | Feb 26, 2026 | ||
| Object Goal Navigation | HM3D | 96 | Jun 17, 2026 | ||
| Graph Classification | Wiki-CS | 96 | Mar 11, 2026 | ||
| Text-to-Image Generation | Qwen-Image |
| 96 |
| May 25, 2026 |
| Outlier detection | ADNI (test) | 96 | Feb 26, 2026 |
|---|
| Long-context Evaluation | LongBench | 96 | Jul 2, 2026 |
|---|
| Information Retrieval | PIRE multi-chunk | 96 | Feb 26, 2026 |
|---|
| Information Retrieval | PIRE single-chunk 1.0 (test) | 96 | Feb 26, 2026 |
|---|
| Text-to-Image Generation | GenEval | 96 | Mar 6, 2026 |
|---|
| Black-box Attack | VLM Evaluation Set (test) | 96 | Feb 26, 2026 |
|---|
| Surface Normal Estimation | NYU v2 | 96 | Jul 8, 2026 |
|---|
| Calibration | BEAR (test) | 96 | Feb 26, 2026 |
|---|
| Long-context Memory Retrieval | LoCoMo | 96 | Jun 12, 2026 |
|---|
| Mathematical Problem Solving | MATH500 | 96 | Jun 23, 2026 |
|---|
| Safety Evaluation | DoNotAnswer Framed | 96 | Feb 26, 2026 |
|---|
| Mathematical Reasoning | AIME 2025 | 96 | Feb 26, 2026 |
|---|
| Object Detection | DroneVehicle (test) | 96 | Jul 7, 2026 |
|---|
| LGT Detection | Fast-DetectGPT XSum (test) | 96 | Feb 26, 2026 |
|---|
| LGT Detection | Fast-DetectGPT PubMed (test) | 96 | Feb 26, 2026 |
|---|
| Text-based Visual Question Answering | TextVQA (VQA^T) | 96 | Mar 24, 2026 |
|---|
| Referring Segmentation | refCOCO (val) | 96 | Jul 8, 2026 |
|---|
| Multimodal Understanding | MMBench Chinese | 96 | Jun 24, 2026 |
|---|
| Base-to-New Generalization | OxfordPets | 96 | Jul 2, 2026 |
|---|
| Text-to-Image Generation | PartiPrompts | 96 | Jun 23, 2026 |
|---|