Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Multimodal Understanding | SEEDBench2 Plus | 138 | May 20, 2026 | ||
| Multimodal Reasoning | MathVerse | 138 | Jun 23, 2026 | ||
| Multiple Choice Question Answering | MMLU-Pro | 138 | Jun 4, 2026 | ||
| Mathematical Reasoning | Minerva |
| 138 |
| Feb 26, 2026 |
| Node classification | Cora | 138 | Jun 11, 2026 |
|---|
| RGB-T tracking | GTOT | 138 | Mar 24, 2026 |
|---|
| Reading Comprehension | DROP | 138 | Jun 30, 2026 |
|---|
| Backdoor Defense | GTSRB (test) | 138 | Jun 15, 2026 |
|---|
| Question Answering | Winogrande (WG) | 138 | Apr 21, 2026 |
|---|
| Few-shot classification | miniImageNet standard (test) | 138 | Feb 26, 2026 |
|---|
| Person Re-identification | DukeMTMC-reID to Market-1501 (test) | 138 | Mar 17, 2026 |
|---|
| Semantic Segmentation | Synthia to Cityscapes (test) | 138 | Feb 26, 2026 |
|---|
| Fine-grained classification | EuroSAT | 138 | Jun 23, 2026 |
|---|
| Polyp Segmentation | ETIS | 138 | Jul 2, 2026 |
|---|
| Object Hallucination Evaluation | POPE | 137 | Jul 9, 2026 |
|---|
| Multimodal Understanding | MMBench | 137 | Jul 1, 2026 |
|---|
| Perplexity | C4 | 137 | May 12, 2026 |
|---|
| Model reduction error estimation | Transport equation trajectory manifold Eq 4.4, Init 4.5 | 137 | Feb 26, 2026 |
|---|
| Image Super-Resolution | B100 | 137 | May 22, 2026 |
|---|
| Reward Modeling | RM-Bench | 137 | Jun 9, 2026 |
|---|
| Massive Multitask Language Understanding | MMLU | 137 | Jul 3, 2026 |
|---|
| Multi-turn Instruction Following | MT Bench | 137 | Jun 11, 2026 |
|---|
| Science Question Answering | SciQ | 137 | Mar 17, 2026 |
|---|
| Unconditional Image Generation | CIFAR-10 32x32 (test) | 137 | Mar 13, 2026 |
|---|
| Image Classification | CIFAR10 | 137 | Mar 20, 2026 |
|---|