Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Chinese Language Understanding | C-Eval | 68 | May 26, 2026 | ||
| Deepfake Detection | Celeb-DF v2 (test) | 68 | Apr 24, 2026 | ||
| Code Generation | CodeContests | 68 | May 18, 2026 | ||
| Few-shot Classification | ModelNet40 (test) | 68 | Feb 26, 2026 | ||
| Symbolic Reasoning | Last Letter Concatenation |
| 68 |
| May 14, 2026 |
| Referring Segmentation | refCOCO+ (val) | 68 | Jul 8, 2026 |
|---|
| Image Classification | CIFAR10 long-tailed (test) | 68 | Feb 26, 2026 |
|---|
| Music Genre Classification | GTZAN | 68 | Jun 18, 2026 |
|---|
| Web agent tasks | Mind2Web Cross-Task | 68 | Jun 9, 2026 |
|---|
| Commonsense Reasoning | HellaSwag (val) | 68 | Jun 23, 2026 |
|---|
| Camera Tracking | Replica | 68 | Jun 30, 2026 |
|---|
| Video Super-Resolution | SPMCS | 68 | Jun 9, 2026 |
|---|
| Unconditional Image Generation | LSUN Bedroom 256x256 | 68 | Mar 23, 2026 |
|---|
| Dynamic Link Prediction | Can. Parl. Inductive | 68 | Jun 4, 2026 |
|---|
| Traffic Forecasting | PEMS03 | 68 | Jun 12, 2026 |
|---|
| General Reasoning | BIG-Bench Hard | 68 | May 26, 2026 |
|---|
| Question Answering | CSQA (test) | 68 | May 15, 2026 |
|---|
| Irregularly Sampled Time Series Forecasting | USHCN (test) | 68 | May 28, 2026 |
|---|
| Image-to-text retrieval | MSCOCO 5K (test) | 68 | Apr 21, 2026 |
|---|
| Hierarchical Text Classification | RCV1 v2 | 68 | Apr 20, 2026 |
|---|
| Video Captioning | MSRVTT | 68 | Mar 25, 2026 |
|---|
| Code Generation | CodeContests (test) | 68 | May 27, 2026 |
|---|
| Offline Reinforcement Learning | D4RL halfcheetah v2 (medium-replay) | 68 | Jun 2, 2026 |
|---|
| Text reconstruction from gradients | Rotten Tomatoes | 68 | Apr 9, 2026 |
|---|
| Image Classification | ImageNet (val) | 68 | Feb 26, 2026 |
|---|