Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Secure Code Generation | Java (evaluation) | 66 | Jun 16, 2026 | ||
| Secure Code Generation | C++ (test) | 66 | Jun 16, 2026 | ||
| Question Answering | NaturalQA | 66 | Jun 11, 2026 | ||
| Question Answering | WebQuestions | 66 | Jun 11, 2026 | ||
| Image Classification | TinyImageNet Second Session Transition |
| 66 |
| Jun 9, 2026 |
| Diagnosis Prediction | Held-out (test) | 66 | Jun 9, 2026 |
|---|
| Face Verification | RFW | 66 | Jun 4, 2026 |
|---|
| Face Verification | LFW | 66 | Jun 4, 2026 |
|---|
| Calibration | LFW | 66 | Jun 4, 2026 |
|---|
| Calibration | BFW | 66 | Jun 4, 2026 |
|---|
| Calibration | RFW | 66 | Jun 4, 2026 |
|---|
| Language Modeling | Language Modeling Scaling Study Dataset | 66 | Jun 4, 2026 |
|---|
| Multiple-Choice Question Answering | World Knowledge Average of OBQA, ARC-C, ARC-E, SCIQ, SIQA | 66 | Jun 2, 2026 |
|---|
| Object Detection | DeepPCB 1.0 (test) | 66 | Jun 2, 2026 |
|---|
| Image Classification | ImageNet-1k | 66 | Jul 7, 2026 |
|---|
| Image Classification | Tiny ImageNet | 66 | Jun 29, 2026 |
|---|
| Reasoning | Reasoning Benchmarks BBH, MMLU, ARC-C, ThmQA (test) | 66 | May 27, 2026 |
|---|
| Multi-label Recognition | NUS-WIDE | 66 | Jun 12, 2026 |
|---|
| Speculative Decoding | LiveCodeBench | 66 | Jun 2, 2026 |
|---|
| Class Erasure | Imagenette | 66 | May 18, 2026 |
|---|
| DRC Script Synthesis | Rule2DRC 1,000 problems | 66 | May 18, 2026 |
|---|
| DRC Script Synthesis | Rule2DRC (test) | 66 | May 18, 2026 |
|---|
| Commonsense Reasoning | HellaSwag (HS) | 66 | Jun 2, 2026 |
|---|
| General Capability | Aggregate (GPQA-D, GSM8K, HumanEval, MATH-500, MBPP, MMLU-Pro) | 66 | May 12, 2026 |
|---|
| Commonsense Reasoning | HellaSwag | 66 | Jun 29, 2026 |
|---|