Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Molecular Optimization (QED) | TOMG-Bench | 39 | May 28, 2026 | ||
| Molecular Optimization (MR) | TOMG-Bench | 39 | May 28, 2026 | ||
| Molecular Optimization (LogP) | TOMG-Bench | 39 | May 28, 2026 | ||
| Multi-modal Long-context Benchmarking | MileBench | 39 | Mar 6, 2026 | ||
| Dynamic Link Prediction | UCI (inductive split) | 39 | Jun 4, 2026 | ||
| Skin Lesion Segmentation |
| HAM10000 |
| 39 |
| Jun 26, 2026 |
| Offline Reinforcement Learning | D4RL Maze2D Large | 39 | Jun 15, 2026 |
|---|
| Video Editing | OpenVE-Bench | 39 | Jun 1, 2026 |
|---|
| Low-light image enhancement | LOL Synthetic v2 (test) | 39 | May 25, 2026 |
|---|
| Language Modeling | C4 | 39 | Mar 4, 2026 |
|---|
| Binary Classification | Interaction Dataset (DI) | 39 | Mar 4, 2026 |
|---|
| 3D wireframe reconstruction | DTU Dataset | 39 | Feb 27, 2026 |
|---|
| Medical Image Classification | CheXpert | 39 | Jun 4, 2026 |
|---|
| Multimodal Perception and Cognition | MME (test) | 39 | Apr 15, 2026 |
|---|
| Multimodal Understanding | MME | 39 | Mar 10, 2026 |
|---|
| Multi-view Video Generation | nuScenes (val) | 39 | Jul 1, 2026 |
|---|
| Code Generation | HumanEval++ | 39 | Apr 28, 2026 |
|---|
| Image Understanding | MMBench CN | 39 | Mar 5, 2026 |
|---|
| Bayesian network structure discovery | Hailfinder | 39 | Apr 7, 2026 |
|---|
| Medical Image Segmentation | BTCV | 39 | Jul 7, 2026 |
|---|
| Mathematical Problem Solving | AIME 24 | 39 | May 12, 2026 |
|---|
| Classification | BBBP | 39 | May 26, 2026 |
|---|
| Language Modeling | Lambada (val) | 39 | May 26, 2026 |
|---|
| Time Series Classification | UEA-27 (test) | 39 | Jul 7, 2026 |
|---|
| Node classification | Wisconsin (Wis.) | 39 | Jun 25, 2026 |
|---|