Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Graph Anomaly Detection | Questions | 59 | May 26, 2026 | ||
| Question Answering | 2Wiki (test) | 59 | Jun 30, 2026 | ||
| Zero-shot evaluation | ARC-Easy, ARC-Challenge, OpenBookQA, WinoGrande, PIQA, HellaSwag, MathQA, RTE, BoolQ zero-shot | 59 | Mar 10, 2026 | ||
| Spatial Reasoning | SPAR-Bench | 59 | Jul 7, 2026 | ||
| Summarization | TL;DR |
| 59 |
| May 27, 2026 |
| Long-context evaluation | RULER | 59 | Jun 19, 2026 |
|---|
| Video Understanding | LongVideoBench | 59 | Jul 1, 2026 |
|---|
| Classification | Churn | 59 | May 15, 2026 |
|---|
| Long-form Generation | Bio | 59 | Apr 2, 2026 |
|---|
| Mathematical Reasoning | AIME 2025 | 59 | Mar 4, 2026 |
|---|
| Readmission Prediction | MIMIC-III (target) | 59 | Jun 5, 2026 |
|---|
| Code Generation | MBPP | 59 | Mar 9, 2026 |
|---|
| Classification | wisconsin | 59 | May 13, 2026 |
|---|
| Binary Classification | Haberman | 59 | Apr 10, 2026 |
|---|
| Language Modeling | WikiText and LAMBADA | 59 | May 8, 2026 |
|---|
| Hallucination Evaluation | HallBench | 59 | Jul 3, 2026 |
|---|
| GUI Grounding | UI-Vision (test) | 59 | May 28, 2026 |
|---|
| 3D semantic occupancy prediction | SemanticKITTI (val) | 59 | Jul 7, 2026 |
|---|
| Document Parsing | olmOCR-bench | 59 | May 4, 2026 |
|---|
| Multimodal Search-based Question Answering | MMSearch | 59 | Jun 30, 2026 |
|---|
| Real-World Image Super-Resolution | RealLQ250 | 59 | May 12, 2026 |
|---|
| 3D Semantic Occupancy Prediction | SurroundOcc-nuScenes (val) | 59 | Apr 2, 2026 |
|---|
| Super-Resolution | RealLQ250 | 59 | Jun 9, 2026 |
|---|
| Robot Manipulation | LIBERO-Plus Zero-shot | 59 | Jul 3, 2026 |
|---|
| Response Harmfulness Detection | BeaverTails | 59 | Jun 1, 2026 |
|---|