Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Model Editing | UltraEditBench | 78 | May 13, 2026 | ||
| Clinical Event Prediction | MIMIC-IV (test) | 78 | Mar 10, 2026 | ||
| LLM Decoding | LLM Decoding | 78 | Mar 6, 2026 | ||
| Coding | MBPP | 78 | May 7, 2026 | ||
| Language Modeling | OWT | 78 | Jun 1, 2026 | ||
| Membership Inference Attack | ELD user-level (test) |
| 78 |
| Feb 26, 2026 |
| Membership Inference Attack | ELD record-level (test) | 78 | Feb 26, 2026 |
|---|
| Membership Inference Attack | ELD | 78 | Feb 26, 2026 |
|---|
| Membership Inference Attack | TUH-EEG | 78 | Feb 26, 2026 |
|---|
| Regression | Communities and Crime 1990 US Census / 1990 US LEMAS / 1995 FBI UCR (test (20%)) | 78 | Feb 26, 2026 |
|---|
| Regression | Criteo Sponsored Search Conversion Log (test) | 78 | Feb 26, 2026 |
|---|
| Regression | California Housing Standard (test) | 78 | Feb 26, 2026 |
|---|
| Regression | California Housing | 78 | Jun 12, 2026 |
|---|
| Multi-Objective Offline Policy Evaluation | MIMIC-IV (test) | 78 | Apr 24, 2026 |
|---|
| Remaining Useful Life prediction | C-MAPSS FD001 | 78 | Jun 18, 2026 |
|---|
| Recognizing Textual Entailment | RTE | 78 | Jun 2, 2026 |
|---|
| Offline Reinforcement Learning | OGBench | 78 | Jul 8, 2026 |
|---|
| Human Activity Recognition | UP-Fall | 78 | Feb 26, 2026 |
|---|
| Human Activity Recognition | DailySport | 78 | Feb 26, 2026 |
|---|
| Subspace Clustering | HSI-Pavia 10 classes | 78 | Feb 26, 2026 |
|---|
| GUI Grounding | ScreenSpot Web V2 | 78 | Jul 1, 2026 |
|---|
| GUI Grounding | ScreenSpot Desktop V2 | 78 | Jul 1, 2026 |
|---|
| GUI Grounding | ScreenSpot Mobile V2 | 78 | Jul 1, 2026 |
|---|
| Video Semantic Segmentation | Cityscapes-C (test) | 78 | Feb 26, 2026 |
|---|
| Hallucination Evaluation | Object HalBench | 78 | Jun 30, 2026 |
|---|