Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Semantic Segmentation | LoveDA (test) | 92 | Apr 30, 2026 | ||
| Reasoning | OpenBookQA | 92 | May 20, 2026 | ||
| Semantic Segmentation | COCO Stuff-27 (val) | 92 | May 28, 2026 | ||
| Role-Playing | Alpaca-P | 91 | Jun 15, 2026 | ||
| SAR Image Classification | MSTAR publicly released | 91 | May 19, 2026 | ||
| Behavior decoding | Makin (held-out sessions) |
| 91 |
| May 15, 2026 |
| Drift detection | Controlled (post-cutoff) | 91 | May 12, 2026 |
|---|
| Question Answering | IntraBench (test) | 91 | Apr 28, 2026 |
|---|
| Hallucination Detection | TruthfulQA | 91 | Jun 2, 2026 |
|---|
| Node Classification | Physics | 91 | Jul 3, 2026 |
|---|
| Jailbreak Defense | HarmBench | 91 | Jun 1, 2026 |
|---|
| Visible-Infrared Person Re-Identification | SYSU-MM01 All-Search | 91 | Jun 11, 2026 |
|---|
| General Language Understanding | tinyBenchmark | 91 | Jun 16, 2026 |
|---|
| Code Generation | BigCodeBench Full | 91 | Jun 18, 2026 |
|---|
| Robotic Manipulation | LIBERO Spatial Object Goal Long | 91 | Jun 23, 2026 |
|---|
| Object Hallucination Evaluation | A-OKVQA POPE (Popular) | 91 | Jun 30, 2026 |
|---|
| Novel-view synthesis | RE10K (test) | 91 | Jun 4, 2026 |
|---|
| Video Understanding | MMVU | 91 | Jul 7, 2026 |
|---|
| Optical Character Recognition Evaluation | OCRBench | 91 | May 21, 2026 |
|---|
| Spatial Reasoning | MindCube | 91 | Jun 11, 2026 |
|---|
| Prompt Injection | MMLU | 91 | May 19, 2026 |
|---|
| General Reasoning | StratQA | 91 | Feb 26, 2026 |
|---|
| Code Generation | LiveCodeBench v6 | 91 | Jun 17, 2026 |
|---|
| Hallucination Detection | NQ (test) | 91 | May 7, 2026 |
|---|
| Autonomous Driving Planning | NAVSIM (navtest) | 91 | Jul 3, 2026 |
|---|