Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Question Answering | TriviaQA | 62 | Feb 26, 2026 | ||
| Sequential Recommendation | Sports | 62 | Feb 26, 2026 | ||
| Graph Anomaly Detection | Cora | 62 | Jun 12, 2026 | ||
| Image Classification | FGVC Aircraft | 62 | Jun 23, 2026 | ||
| Remote Sensing Classification | SIRI-WHU | 62 | Jun 23, 2026 | ||
| Inpainting | FFHQ |
| 62 |
| May 12, 2026 |
| Emotion Recognition | MER-UniBench (test) | 62 | Jun 16, 2026 |
|---|
| Zero-shot Classification | CIFAR10 | 62 | May 20, 2026 |
|---|
| Multimodal Benchmarking | MMBench (MMB) | 62 | Mar 10, 2026 |
|---|
| Image Editing | ImgEdit | 62 | May 21, 2026 |
|---|
| Website Fingerprinting | H&W-1600 (test) | 62 | Feb 26, 2026 |
|---|
| Tool-use | ToolBench | 62 | May 27, 2026 |
|---|
| Commonsense Reasoning | CommonsenseQA (CSQA) | 62 | Jun 17, 2026 |
|---|
| Code Question Answering | CodeSimpleQA Chinese | 62 | Feb 26, 2026 |
|---|
| Question Answering | Musique | 62 | Apr 21, 2026 |
|---|
| Logical Reasoning | HLE | 62 | Apr 23, 2026 |
|---|
| Long-context language understanding | LongBench v2 | 62 | May 13, 2026 |
|---|
| Mathematical Reasoning | Math Benchmarks Aggregate | 62 | May 28, 2026 |
|---|
| Long-context evaluation | RULER 16k | 62 | Jun 9, 2026 |
|---|
| Reward Modeling | HelpSteer 3 | 62 | May 29, 2026 |
|---|
| Jailbreak Attack | JailbreakBench (JBB) | 62 | Apr 10, 2026 |
|---|
| Hallucination Detection | CommonsenseQA | 62 | Apr 20, 2026 |
|---|
| in-hospital mortality prediction | MIMIC-IV | 62 | Apr 21, 2026 |
|---|
| Mobile GUI Automation | GUI-Odyssey | 62 | Apr 8, 2026 |
|---|
| Semantic Grounding | COCO 2017 (val) | 62 | Feb 26, 2026 |
|---|