Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Portfolio Optimization | Multivariate Return Series clean data (evaluation) | 35 | Feb 26, 2026 | ||
| Multimodal Search | BrowseComp-VL | 35 | Jul 1, 2026 | ||
| Scientific Question Answering | SciQA | 35 | May 26, 2026 | ||
| Language Generation | CNN/DailyMail | 35 | Feb 26, 2026 | ||
| Language Generation | XSum | 35 | Feb 26, 2026 | ||
| Language Generation |
| QASPER |
| 35 |
| Feb 26, 2026 |
| Language Generation | CoQA | 35 | Feb 26, 2026 |
|---|
| Reranking | TREC | 35 | Feb 26, 2026 |
|---|
| Reranking | BEIR | 35 | Feb 26, 2026 |
|---|
| Multimodal Embedding Evaluation | MMEB V2 (test) | 35 | Apr 8, 2026 |
|---|
| Mathematical Reasoning | AIME 24 | 35 | Feb 27, 2026 |
|---|
| Camera pose estimation | Oxford Spires | 35 | Jun 4, 2026 |
|---|
| Intrusion Detection | NSL-KDD (test) | 35 | Jun 12, 2026 |
|---|
| Multi-Agent Reinforcement Learning | SMAC v2 (test) | 35 | Jun 25, 2026 |
|---|
| Language Modeling | C4 | 35 | Feb 26, 2026 |
|---|
| Regression | California Housing (test) | 35 | Jun 23, 2026 |
|---|
| Automatic Speech Recognition | LJ-Speech | 35 | Feb 26, 2026 |
|---|
| Automatic Speech Recognition | LibriSpeech | 35 | Feb 26, 2026 |
|---|
| Speech Recognition | LJ-Speech (test) | 35 | Feb 26, 2026 |
|---|
| Voice Activity Detection | LibriVAD-Concat (small-scale test) | 35 | Feb 26, 2026 |
|---|
| Voice Activity Detection | LibriVAD-NonConcat small-scale (train and test) | 35 | Feb 26, 2026 |
|---|
| Robotic Manipulation | SIMPLER Google Robot VA | 35 | Mar 27, 2026 |
|---|
| Long Video Generation | VBench | 35 | May 27, 2026 |
|---|
| Closed-loop Planning | nuPlan random 14 (test) | 35 | Jun 30, 2026 |
|---|
| Robot Manipulation | LIBERO (All four suites (combined)) | 35 | Jun 5, 2026 |
|---|