Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Brain Tumor Segmentation | BraTS T1–FLAIR (train) | 57 | Feb 26, 2026 | ||
| Novel-view synthesis | RE10K (Average) | 57 | Jul 1, 2026 | ||
| Novel-view synthesis | RE10K (Medium) | 57 | Jul 1, 2026 | ||
| Image Classification | CUB200 | 57 | Jun 16, 2026 | ||
| Instruction Following | ALFRED | 57 | May 7, 2026 | ||
| Object Hallucination Assessment |
| A-OKVQA POPE (Adversarial) |
| 57 |
| Jun 30, 2026 |
| Visual Question Answering | FVQA | 57 | Jul 1, 2026 |
|---|
| Synthetic Image Detection | DRCT-2M | 57 | May 22, 2026 |
|---|
| Hallucination Detection | VQA-RAD (All) | 57 | May 21, 2026 |
|---|
| Hallucination Detection | VQA-RAD Open-Ended | 57 | May 21, 2026 |
|---|
| Multi-Modal Visual Question Answering (MMVQA) | RAD-ChestCT (val) | 57 | Feb 26, 2026 |
|---|
| Multi-Modal Visual Question Answering (MMVQA) | CT-RATE (val) | 57 | Feb 26, 2026 |
|---|
| Out-of-Distribution Detection | CIFAR-10 In-Dist Texture Out-Dist | 57 | Apr 28, 2026 |
|---|
| Few-shot Semantic Segmentation | Average Deepglobe, ISIC, Chest X-Ray, FSS-1000 | 57 | Jun 24, 2026 |
|---|
| Science Reasoning | GPQA Diamond | 57 | Jun 11, 2026 |
|---|
| Mathematical Reasoning | MATH500 1.0 (test) | 57 | May 13, 2026 |
|---|
| LLM Inference | Alpaca | 57 | Jun 9, 2026 |
|---|
| Mathematical Reasoning | IMO-Bench | 57 | May 22, 2026 |
|---|
| Aggregate Model Performance | Combined Benchmark Suite | 57 | Mar 27, 2026 |
|---|
| General Reasoning | GPQA Diamond | 57 | Jun 9, 2026 |
|---|
| Audio Understanding | MMAU | 57 | Jun 30, 2026 |
|---|
| Error Detection | HotpotQA | 57 | Apr 23, 2026 |
|---|
| Math | AIME24 | 57 | Jun 2, 2026 |
|---|
| General Reasoning | MMMU | 57 | Jul 3, 2026 |
|---|
| Multi-modal Video Evaluation | Video-MME | 57 | Jun 16, 2026 |
|---|