Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Image Classification | ImageNet source to 10 fine-grained target datasets (test) | 37 | May 12, 2026 | ||
| Temporal Forgery Localization | LAV-DF (test) | 37 | Jul 7, 2026 | ||
| General Evaluation | Aggregate Benchmarks | 37 | May 13, 2026 | ||
| User Badge Prediction | Rel Stack User Badge | 37 | May 25, 2026 | ||
| Multimodal Perception Evaluation | MME-P | 37 |
| May 27, 2026 |
| Autonomous Driving Planning | NAVSIM v2 | 37 | Jun 30, 2026 |
|---|
| Knowledge | GPQA 0-shot | 37 | Mar 4, 2026 |
|---|
| Knowledge | MMLU-Pro 5-shot | 37 | Mar 4, 2026 |
|---|
| Multimodal Retrieval | MMEB Image V2 | 37 | Jun 11, 2026 |
|---|
| Code Generation | LiveCodeBench (LCB) | 37 | Jun 17, 2026 |
|---|
| Node classification | Walmart | 37 | May 12, 2026 |
|---|
| Truthful Question Answering | TruthfulQA | 37 | Jun 24, 2026 |
|---|
| Survival Prediction | TCGA GBM-LGG Internal (test) | 37 | Mar 12, 2026 |
|---|
| Numerical Optimization | CEC 50 dimensions 2017 | 37 | May 12, 2026 |
|---|
| Reasoning | BigBenchHard | 37 | Jul 7, 2026 |
|---|
| Unified Multi-task Language Understanding and Instruction Following | Open LLM Leaderboard v1 (test) | 37 | Jun 5, 2026 |
|---|
| Image Classification | CIFAR100 | 37 | Apr 21, 2026 |
|---|
| Imputation | Electricity | 37 | Mar 13, 2026 |
|---|
| Medical Image Segmentation | MoNuSeg | 37 | Jul 3, 2026 |
|---|
| Image Classification | CIFAR-100 Dir-0.5 | 37 | Apr 21, 2026 |
|---|
| Aspect-based Sentiment Analysis | REST 2014 (test) | 37 | May 21, 2026 |
|---|
| Face Forgery Detection | FaceShifter HQ (FSh) | 37 | Mar 4, 2026 |
|---|
| Unconditional image generation | CelebA-HQ 256x256 | 37 | May 18, 2026 |
|---|
| Interactive Decision Making | WebShop (test) | 37 | May 19, 2026 |
|---|
| Super-Resolution | RASMD | 37 | Feb 26, 2026 |
|---|