Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Semantic Segmentation | NYU v2 (Retained set) | 37 | May 20, 2026 | ||
| Multi-turn embodied reasoning | BabyAI | 37 | May 26, 2026 | ||
| Zero-shot Segmentation | MS-COCO | 37 | May 19, 2026 | ||
| Single Image Reflection Removal | Postcard 199 (test) | 37 | Jul 1, 2026 | ||
| Multi-objective Optimization | MKP | 37 | May 18, 2026 | ||
| Multi-objective Optimization |
| MMMP |
| 37 |
| May 18, 2026 |
| Single-Hop Question Answering | NQ | 37 | May 14, 2026 |
|---|
| 3D Shape Correspondence | SCAPE_r (test) | 37 | May 14, 2026 |
|---|
| 3D Shape Correspondence | FAUST_r (test) | 37 | May 14, 2026 |
|---|
| Text-to-Image Generation | OneIG-EN (test) | 37 | May 26, 2026 |
|---|
| UI Agent Interaction | AgentNetBench | 37 | Jun 15, 2026 |
|---|
| Object Detection | FLIR Aligned | 37 | Jul 7, 2026 |
|---|
| Robot Manipulation | RoboTwin 2.0 | 37 | Jul 1, 2026 |
|---|
| Reasoning | GPQA | 37 | May 20, 2026 |
|---|
| Perception and Reasoning | RealWorldQA | 37 | Jun 16, 2026 |
|---|
| Text Classification | Civil Comments (test) | 37 | May 12, 2026 |
|---|
| Long-context language understanding | LongBench | 37 | Jun 30, 2026 |
|---|
| Text Question Answering | LongBench | 37 | May 11, 2026 |
|---|
| Text Question Answering | Qasper | 37 | May 11, 2026 |
|---|
| Text Question Answering | RULER | 37 | May 11, 2026 |
|---|
| Text Question Answering | HELMET | 37 | May 11, 2026 |
|---|
| Text Question Answering | MuSiQue | 37 | May 11, 2026 |
|---|
| Text Question Answering | NQ | 37 | May 11, 2026 |
|---|
| Zero-shot Accuracy | Zero-shot Evaluation Suite (PIQA, HellaSwag, MMLU, HumanEval, BoolQ, WinoGrande, ARC-E, ARC-C) (test) | 37 | May 8, 2026 |
|---|
| Coding | MBPP | 37 | Jun 2, 2026 |
|---|