Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Multiple Choice Question Answering | SciQ | 91 | May 25, 2026 | ||
| Speech Recognition | Hub5'00 SWB (test) | 91 | Feb 26, 2026 | ||
| Graph Classification | ENZYMES (test) | 91 | Jul 7, 2026 | ||
| Graph Classification | IMDB-M (10-fold cross-validation) | 91 | May 5, 2026 | ||
| Image Quality Assessment | KonIQ-10k (test) |
| 91 |
| Feb 26, 2026 |
| Person Re-identification | LTCC General | 91 | Jun 11, 2026 |
|---|
| Named Entity Recognition | Wnut 2017 | 91 | Apr 23, 2026 |
|---|
| LoRA Routing | PaperQA PMD-BENCH | 90 | Jul 7, 2026 |
|---|
| Top-10 Attended-Token Overlap | Attention Diagnostics | 90 | Jun 24, 2026 |
|---|
| Paraphrase Detection | MRPC | 90 | Jun 15, 2026 |
|---|
| One-vs-all classification | CIFAR-10 | 90 | Jun 11, 2026 |
|---|
| One-year mortality prediction | MIMIC-IV | 90 | Jun 11, 2026 |
|---|
| Reconstruction | CDC Diabetes 1k rows QIdemo | 90 | Jun 9, 2026 |
|---|
| Robot Manipulation | LIBERO | 90 | Jun 17, 2026 |
|---|
| Error Prediction | AmbigQA (val) | 90 | Jun 2, 2026 |
|---|
| Logical Reasoning | ZebraLogic v1.0 (test) | 90 | May 29, 2026 |
|---|
| topic membership estimation | trimmed ICLR corpus | 90 | May 28, 2026 |
|---|
| Long-context Understanding | RULER 16k (test) | 90 | May 26, 2026 |
|---|
| Long-context Understanding | RULER 4k (test) | 90 | May 26, 2026 |
|---|
| Reward Modeling | RewardBench 2 | 90 | Jun 9, 2026 |
|---|
| Jailbreaking | AdvBench selected models | 90 | May 25, 2026 |
|---|
| Membership Inference Attack | NYU v2 | 90 | May 20, 2026 |
|---|
| 3D Object Detection | MultiCorrupt nuScenes (val) | 90 | May 13, 2026 |
|---|
| Single-hop Tool Calling | WHEN2TOOL single-hop 1.0 (test) | 90 | May 12, 2026 |
|---|
| kNN Search | ModelNet40 | 90 | Apr 24, 2026 |
|---|