Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Hallucination Evaluation | CHAIR MSCOCO 2014 (val) | 45 | May 20, 2026 | ||
| Multi-modal Reasoning | MathVision (test) | 45 | Mar 24, 2026 | ||
| Visual Mathematical Reasoning | DynaMath | 45 | Mar 12, 2026 | ||
| Vulnerability Evasion | CWE-based Vulnerabilities | 45 | Feb 26, 2026 | ||
| Math Word Problem Solving | SOMADHAN (test) |
| 45 |
| Feb 26, 2026 |
| Tool Use | BFCL | 45 | Jun 2, 2026 |
|---|
| Regression | FreeSolv | 45 | May 26, 2026 |
|---|
| Mathematical Reasoning | Beyond AIME | 45 | Mar 18, 2026 |
|---|
| Long-context Question Answering | LoCoMo | 45 | May 12, 2026 |
|---|
| Temporal Reasoning | LoCoMo | 45 | Apr 15, 2026 |
|---|
| Complex Tabular Reasoning | TableBench | 45 | Jun 2, 2026 |
|---|
| Multilingual Information Retrieval | Belebele | 45 | Jun 18, 2026 |
|---|
| Reward Modeling | Aggregate of 7 benchmarks (HelpSteer3, Reward Bench V2, SCAN-HPD, HREF, LitBench, WQ_Arena, WPB) | 45 | Feb 26, 2026 |
|---|
| Logical Fallacy Detection | LOGICAL FALLACY (LOG) | 45 | Jun 26, 2026 |
|---|
| Question Answering | SQuAD (test) | 45 | Feb 26, 2026 |
|---|
| Toxicity Detection | ToxicChat | 45 | Jun 1, 2026 |
|---|
| Rebuttal Quality Evaluation | R2 (test) | 45 | Feb 26, 2026 |
|---|
| Medical Knowledge Question Answering | Medical Domain (MedQA, MMLU, MedMCQA) (test) | 45 | Feb 26, 2026 |
|---|
| Language Modeling | AG News | 45 | Jun 30, 2026 |
|---|
| Mathematical Reasoning | IMO-AnswerBench | 45 | Jul 9, 2026 |
|---|
| Text-to-SQL | Archer (dev) | 45 | Jun 11, 2026 |
|---|
| Code Generation | HumanEval | 45 | Feb 26, 2026 |
|---|
| Model Performance Prediction | DeepSeek Model Families (Hold-out) | 45 | Feb 26, 2026 |
|---|
| Reward Modeling | PPE Correctness | 45 | Jun 1, 2026 |
|---|
| Retrieval | Musique | 45 | Apr 21, 2026 |
|---|