Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Topic Modeling | TeslaModel3 | 44 | Apr 15, 2026 | ||
| Topic Modeling | Bothering | 44 | Apr 15, 2026 | ||
| Visual Question Answering | TextVQA | 44 | Jun 26, 2026 | ||
| Mathematical Reasoning | DeepMath | 44 | Apr 15, 2026 | ||
| Backdoor Defense | MNLI (test) | 44 | Apr 14, 2026 | ||
| Text Classification |
| AG NEWS RoBERTa-large (test) |
| 44 |
| Apr 14, 2026 |
| Personalized Image Aesthetic Assessment | PARA | 44 | Apr 21, 2026 |
|---|
| Image Classification | Stanford Cars | 44 | May 26, 2026 |
|---|
| Document Parsing | OmniDocBench Full v1.6 | 44 | May 28, 2026 |
|---|
| Reference-free Conversation Evaluation | CRSArena-Eval KG | 44 | Apr 7, 2026 |
|---|
| Reference-free Conversation Evaluation | CRSArena-Eval RD | 44 | Apr 7, 2026 |
|---|
| Document Classification | 20 Newsgroups | 44 | Jun 2, 2026 |
|---|
| Offline Synthetic Data Generation | Alpaca | 44 | Apr 2, 2026 |
|---|
| Robot Manipulation | LIBERO LONG | 44 | Jul 7, 2026 |
|---|
| Text-to-Image Retrieval | NWPU (test) | 44 | Mar 31, 2026 |
|---|
| Image-to-Text Retrieval | NWPU (test) | 44 | Mar 31, 2026 |
|---|
| HD map construction | nuScenes (val) | 44 | Jul 1, 2026 |
|---|
| Emotion Recognition | SEED, SEED-IV, and SEED-V | 44 | Mar 31, 2026 |
|---|
| VIO Initialization | EuRoC (All sequences) | 44 | Mar 30, 2026 |
|---|
| Video Temporal Grounding | QVHighlights | 44 | Jun 30, 2026 |
|---|
| Multimodal Reward Modeling | RewardBench Multimodal | 44 | May 12, 2026 |
|---|
| Theory of Mind reasoning | MMToM-QA | 44 | May 28, 2026 |
|---|
| Node Classification | GOODCora Covariate shift, degree (test) | 44 | Mar 26, 2026 |
|---|
| Mathematical Reasoning | AMC | 44 | May 14, 2026 |
|---|
| Digital Watermarking | CelebA (test) | 44 | May 12, 2026 |
|---|