Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Commonsense Reasoning | HellaSwag | 47 | Apr 10, 2026 | ||
| Test-Time Adaptation | ImageNet-C level 5 (5k) | 47 | Jun 2, 2026 | ||
| Sentiment Analysis | SST2 | 47 | May 6, 2026 | ||
| Object Classification | ScanObjectNN PB-T50-RS | 47 | Jul 8, 2026 | ||
| Anomaly Detection | MVTec 3D-AD | 47 | Mar 10, 2026 | ||
| Common-sense Reasoning |
| Common-sense Reasoning Average |
| 47 |
| Jul 2, 2026 |
| Single-Hop | LoCoMo | 47 | May 14, 2026 |
|---|
| Medical Image Classification | OCTMNIST | 47 | May 27, 2026 |
|---|
| Offline Reinforcement Learning | D4RL MuJoCo walker2d-medium-expert | 47 | Jun 11, 2026 |
|---|
| Offline Reinforcement Learning | D4RL MuJoCo hopper-medium-expert | 47 | Jun 11, 2026 |
|---|
| Offline Reinforcement Learning | D4RL MuJoCo halfcheetah-medium-replay | 47 | Jun 11, 2026 |
|---|
| Classification | blood | 47 | Jun 11, 2026 |
|---|
| Language Model Pre-training | C4 Llama 2 pre-training (val) | 47 | Feb 26, 2026 |
|---|
| Multi-task Knowledge Understanding | MMLU-Pro | 47 | Jun 25, 2026 |
|---|
| Language Model Evaluation | Open LLM Leaderboard v2 (test) | 47 | May 12, 2026 |
|---|
| Temporal Video Grounding | Charades-STA | 47 | Apr 10, 2026 |
|---|
| Temporal | LoCoMo | 47 | May 14, 2026 |
|---|
| Temporal Knowledge Graph Forecasting | ICEWS 18 | 47 | Apr 8, 2026 |
|---|
| Classification | Fashion MNIST | 47 | May 6, 2026 |
|---|
| Keyword Spotting | SHD (test) | 47 | Feb 26, 2026 |
|---|
| Language Modeling | OpenWebText (OWT) (val) | 47 | Jul 3, 2026 |
|---|
| Robotic Manipulation | Robomimic Lift | 47 | Jun 24, 2026 |
|---|
| Natural Language Understanding | GLUE (test) | 47 | Jun 2, 2026 |
|---|
| Image Classification | SVHN | 47 | May 12, 2026 |
|---|
| Reinforcement Learning | Halfcheetah v5 | 47 | May 1, 2026 |
|---|