Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Counterfactual Generation | SNLI Hypothesis | 37 | Feb 26, 2026 | ||
| Counterfactual Generation | SNLI Premise | 37 | Feb 26, 2026 | ||
| Counterfactual Generation | AG News | 37 | Feb 26, 2026 | ||
| Counterfactual Generation | IMDb | 37 | Feb 26, 2026 | ||
| Data Compression | enwik9 1GB (test) |
| 37 |
| Feb 26, 2026 |
| Multimodal Retrieval and Understanding | MMEB V2 (test) | 37 | May 19, 2026 |
|---|
| Multi-hop Question Answering | Multi-Hop QA | 37 | Apr 24, 2026 |
|---|
| Long-context language modeling | LongBench-E 1.0 (test) | 37 | Feb 26, 2026 |
|---|
| Audiovisual Video Captioning | SALMONN 2 (test) | 37 | Mar 17, 2026 |
|---|
| Readmission Prediction (RA) | MIMIC-IV (test) | 37 | Apr 27, 2026 |
|---|
| Mathematical reasoning | GSM8K Platinum | 37 | Feb 26, 2026 |
|---|
| Transferable Adversarial Attack | HarmBench Classifier (test) | 37 | Feb 26, 2026 |
|---|
| Math Reasoning | AIME 2024 | 37 | Feb 26, 2026 |
|---|
| Visual Question Answering | countbenchqa | 37 | May 29, 2026 |
|---|
| Knowledge Tracing | Junyi | 37 | May 12, 2026 |
|---|
| Text2SQL | Spider (test) | 37 | Feb 26, 2026 |
|---|
| Long-context reasoning and retrieval | LoCoMo (test) | 37 | Feb 26, 2026 |
|---|
| Hallucination Robustness | HallusionBench | 37 | Jun 19, 2026 |
|---|
| LLM Routing | MMR-Bench | 37 | May 13, 2026 |
|---|
| 4D occupancy forecasting | Occ3D-nuScenes | 37 | Jun 5, 2026 |
|---|
| Image Captioning | RSICD | 37 | Mar 11, 2026 |
|---|
| Retrieval-Augmented Generation | ICR2 | 37 | Feb 26, 2026 |
|---|
| Open-Domain Question Answering | WQ (test) | 37 | Feb 26, 2026 |
|---|
| Radiology Report Generation | CHEXPERT Plus | 37 | Mar 17, 2026 |
|---|
| Low-Light Image Enhancement | SDSD indoor | 37 | May 14, 2026 |
|---|