Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Image Classification | CIFAR-10 | 56 | Feb 26, 2026 | ||
| Subgraph Reconstruction Attack | HCM | 56 | Feb 27, 2026 | ||
| Subgraph Reconstruction Attack | Enron | 56 | Feb 27, 2026 | ||
| Formal Theorem Proving | PutnamBench | 56 | Jun 12, 2026 | ||
| Natural Language Inference | BioNLI | 56 | Feb 26, 2026 | ||
| Multilingual Multiple-Choice Question Answering |
| HeadQA 1.0 (test) |
| 56 |
| Feb 26, 2026 |
| Toxicity Detection | Toxicity Detection 64 model-persona combinations (8 models x 8 personas) | 56 | Feb 26, 2026 |
|---|
| Code Reasoning | CRUXEval | 56 | Mar 18, 2026 |
|---|
| Factual Knowledge Evaluation | PopQA | 56 | May 29, 2026 |
|---|
| Zero-shot reasoning | ZeroShot 7 | 56 | May 25, 2026 |
|---|
| General Knowledge | MMLU-Pro | 56 | Jul 7, 2026 |
|---|
| Planning | Room Domain | 56 | Feb 26, 2026 |
|---|
| Question Answering | MMLU-Pro Natural Setting (test) | 56 | Feb 26, 2026 |
|---|
| Closed-loop planning | nuPlan 14 (test) | 56 | Jul 7, 2026 |
|---|
| Gaze target estimation | VideoAttentionTarget | 56 | Jul 7, 2026 |
|---|
| Question Answering | TriviaQA (TQA) | 56 | May 11, 2026 |
|---|
| Safety Evaluation | HEX-PHI (test) | 56 | Apr 24, 2026 |
|---|
| Hallucination Assessment | AMBER | 56 | Jun 1, 2026 |
|---|
| Human-Object Interaction Detection | HICO-DET (NF-UC) | 56 | Apr 3, 2026 |
|---|
| Articulated Object Reconstruction and Motion Estimation | PARIS Simulation | 56 | Apr 10, 2026 |
|---|
| 3D Medical Image Segmentation | MSWAL | 56 | Feb 26, 2026 |
|---|
| Keypoint Detection | ShapeNetCore V2 (test) | 56 | Feb 26, 2026 |
|---|
| Video Question Answering | MSVD-QA zero-shot (test) | 56 | Feb 26, 2026 |
|---|
| Abstract Reasoning | AbsR | 56 | Feb 26, 2026 |
|---|
| Function-level Code Generation | MBPP+ augmented (test) | 56 | Mar 13, 2026 |
|---|