Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Label Classification | TRDL21 (Experiment 2) | 36 | Feb 26, 2026 | ||
| Reward Modeling | LitBench | 36 | Apr 8, 2026 | ||
| Large Language Model Evaluation | HuggingFace Open LLM Leaderboard lm-eval-harness default (various) | 36 | May 7, 2026 | ||
| Tweet Paraphrasing/Generation | LaMP Tweet | 36 | Apr 27, 2026 | ||
| News Headline Generation | LaMP News | 36 | Apr 27, 2026 | ||
| Hierarchical classification |
| AfroScope-Data confusable (test) |
| 36 |
| Feb 26, 2026 |
| Question Answering | HotpotQA | 36 | Feb 26, 2026 |
|---|
| AI-generated text detection | AcademicResearch | 36 | Feb 26, 2026 |
|---|
| Speech-to-Speech Question-Answering | WebQ | 36 | Jul 8, 2026 |
|---|
| Zero-shot Reasoning | ARC-e, Winogrande, HellaSwag, PIQA | 36 | Feb 26, 2026 |
|---|
| Long-context understanding | LongBench V1 | 36 | May 11, 2026 |
|---|
| Error Detection | MuSiQue (val) | 36 | Feb 26, 2026 |
|---|
| Error Detection | Mintaka (val) | 36 | Feb 26, 2026 |
|---|
| Error Detection | HotpotQA (val) | 36 | Feb 26, 2026 |
|---|
| Error Detection | FRAMES (test) | 36 | Feb 26, 2026 |
|---|
| Error Detection | CRAG multi-hop subset (train) | 36 | Feb 26, 2026 |
|---|
| Error Detection | Bamboogle Full | 36 | Feb 26, 2026 |
|---|
| Error Detection | MuSiQue | 36 | Feb 26, 2026 |
|---|
| Error Detection | Mintaka | 36 | Feb 26, 2026 |
|---|
| Error Detection | FRAMES | 36 | Feb 26, 2026 |
|---|
| Error Detection | CRAG | 36 | Feb 26, 2026 |
|---|
| Error Detection | Bamboogle | 36 | Feb 26, 2026 |
|---|
| Safety Evaluation | CocoNot | 36 | Feb 26, 2026 |
|---|
| Medical Reasoning | HealthBench | 36 | May 19, 2026 |
|---|
| Mathematical Reasoning | OlympiadBench | 36 | Feb 26, 2026 |
|---|