Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Question Answering | NaturalQA (test) | 39 | Jul 8, 2026 | ||
| Link Prediction | LastFM (inductive) | 39 | May 26, 2026 | ||
| Hallucination Detection | HaluEvalQA | 39 | Apr 20, 2026 | ||
| Dynamic Facial Expression Recognition | DFEW (test) | 39 | Jul 1, 2026 | ||
| Visual Place Recognition | hk_mapi | 39 | Apr 14, 2026 | ||
| Visual Place Recognition | sf_mapi |
| 39 |
| Apr 14, 2026 |
| Visual Place Recognition | sf v1 | 39 | Apr 14, 2026 |
|---|
| Vulnerability Detection | MoCQ (test) | 39 | Apr 14, 2026 |
|---|
| Text Generation | LM1B | 39 | Jul 7, 2026 |
|---|
| Partial Multi-Label Learning | yeast | 39 | Apr 13, 2026 |
|---|
| Geospatial Reasoning | GeoMMBench (val) | 39 | Apr 13, 2026 |
|---|
| Document Reranking | TREC DL 19 | 39 | May 15, 2026 |
|---|
| Federated Class-Incremental Learning | CIFAR-100 Distribution-based label imbalance | 39 | Apr 13, 2026 |
|---|
| Classification | Cele+Asv leave-one-modality-out (val) | 39 | Apr 10, 2026 |
|---|
| Classification | Fakeavcele leave-one-modality-out (cross-val) | 39 | Apr 10, 2026 |
|---|
| Classification | LAV-DF leave-one-modality-out cross-validation | 39 | Apr 10, 2026 |
|---|
| Zero-shot Common Sense Reasoning | BoolQ | 39 | Jun 4, 2026 |
|---|
| Math Reasoning | AIME 2024 | 39 | May 18, 2026 |
|---|
| Relative Pose Estimation | ScanNet 1500 | 39 | Jul 7, 2026 |
|---|
| Image Quality Assessment Correlation | Mip-NeRF 360 | 39 | Apr 7, 2026 |
|---|
| General Capability Evaluation | General Capability Suite MMLU, GSM8K, HumanEval, IFEval | 39 | Jun 2, 2026 |
|---|
| Closed-loop planning | NAVSIM v1 (navtest) | 39 | Jun 5, 2026 |
|---|
| General Knowledge | MMLU | 39 | Jun 18, 2026 |
|---|
| Web research | BrowseComp zh | 39 | Apr 6, 2026 |
|---|
| Mathematical Reasoning | AIME '24 | 39 | Jun 26, 2026 |
|---|