Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Object Navigation | Navigation Environments (test) | 42 | Jun 11, 2026 | ||
| Mathematical Reasoning | Minerva Math | 42 | Jun 9, 2026 | ||
| Mathematical Reasoning | MATH-OAI | 42 | Jun 9, 2026 | ||
| Language Generation | Language Modeling Evaluation Set | 42 | Jun 9, 2026 | ||
| Mathematical reasoning | AMC23 |
| 42 |
| Jun 11, 2026 |
| Image Classification | CLIP-ViT TA8 | 42 | Jun 8, 2026 |
|---|
| Question Answering | FairytaleQA (test) | 42 | Jun 5, 2026 |
|---|
| Question Answering | STAGE Main Results | 42 | Jun 5, 2026 |
|---|
| Class-Incremental Learning | CUB-200 | 42 | Jun 5, 2026 |
|---|
| 5-way few-shot classification | Dogs | 42 | Jun 5, 2026 |
|---|
| Image Representation | peppers | 42 | Jun 5, 2026 |
|---|
| Image Representation | Lena | 42 | Jun 5, 2026 |
|---|
| LLM Decoding | Long Context 64K | 42 | Jun 4, 2026 |
|---|
| Tactical Move Prediction | Lichess tactical puzzles v18 (reservoir sample (n=74,424)) | 42 | Jun 4, 2026 |
|---|
| Multi-view Clustering | HandWritten (test) | 42 | Jun 4, 2026 |
|---|
| Multi-view Clustering | CCV20 (test) | 42 | Jun 4, 2026 |
|---|
| Multi-view Clustering | Reuters (test) | 42 | Jun 4, 2026 |
|---|
| Multi-view Clustering | Scene15 (test) | 42 | Jun 4, 2026 |
|---|
| Jailbreaking | JailbreakBench | 42 | Jun 4, 2026 |
|---|
| Jailbreak Attack | AdvBench | 42 | Jun 4, 2026 |
|---|
| Image Captioning | COCO-Cap | 42 | Jun 4, 2026 |
|---|
| Question Answering | MuSiQue (held-out) | 42 | Jun 2, 2026 |
|---|
| Medical Image Segmentation | CLI (test) | 42 | Jun 2, 2026 |
|---|
| Medical Image Segmentation | COL (test) | 42 | Jun 2, 2026 |
|---|
| Medical Image Segmentation | DMF (test) | 42 | Jun 2, 2026 |
|---|