Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Panoptic video scene graph generation | OpenPVSG v1 (test) | 48 | Apr 28, 2026 | ||
| Boundary Detection | NYUD v2 | 48 | Jun 16, 2026 | ||
| Vision-Language Model Editing | FVQA 1.0 (test) | 48 | Feb 26, 2026 | ||
| VLM Editing | A-OKVQA 2022 (test) | 48 | Feb 26, 2026 | ||
| Motion-to-Text | HumanML3D (test) |
| 48 |
| Apr 24, 2026 |
| Video Understanding | VideoMME, EgoSchema, LongVideoBench, MVBench | 48 | Feb 27, 2026 |
|---|
| Image Generation | ImageNet-1K 256x256 | 48 | Jul 3, 2026 |
|---|
| Classification | Roadside LiDAR dataset (test) | 48 | Feb 26, 2026 |
|---|
| Image Classification | CIFAR-100 | 48 | Feb 26, 2026 |
|---|
| Image Classification | CIFAR-10 | 48 | Feb 26, 2026 |
|---|
| Imprint Forgery Attack | SDP prompt v1 (val) | 48 | Feb 26, 2026 |
|---|
| Top-k popular visited locations selection | Foursquare check-in Top-3 popular visited locations | 48 | Feb 26, 2026 |
|---|
| Multiple Choice Question Answering | ARC Challenge | 48 | Apr 27, 2026 |
|---|
| Code Generation | BigCodeBench-Instruct Hard | 48 | Mar 16, 2026 |
|---|
| Code Generation | BigCodeBench-Instruct (Full) | 48 | Mar 16, 2026 |
|---|
| Safety Classification | SafeRLHF | 48 | Apr 21, 2026 |
|---|
| Knowledge Reasoning | GPQA | 48 | Jun 5, 2026 |
|---|
| Code Generation | FullStackBench | 48 | May 27, 2026 |
|---|
| Mathematical Reasoning | AIME | 48 | Feb 26, 2026 |
|---|
| Mathematical Reasoning | MATH-500 | 48 | Feb 26, 2026 |
|---|
| Mathematical Reasoning | GSM8K | 48 | Feb 26, 2026 |
|---|
| Jailbreak Attack | ShadowRisk | 48 | Feb 26, 2026 |
|---|
| Jailbreak Attack | AdvBench 50 | 48 | Jul 7, 2026 |
|---|
| Text-to-SQL | Science Benchmark | 48 | Jun 15, 2026 |
|---|
| Dialogue Generation | CONVAI2 | 48 | May 21, 2026 |
|---|