Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Object Detection | BDD100K Dawn Dusk | 40 | May 14, 2026 | ||
| Interactive Decision Making | InterCode NL2Bash | 40 | Apr 21, 2026 | ||
| Interactive Decision Making | WebShop (Seen) | 40 | Apr 21, 2026 | ||
| Question Answering | PIQA (test) | 40 | May 12, 2026 | ||
| Embodied Agent Task | ALFWorld Unseen | 40 | Jun 1, 2026 | ||
| Doc-String Prediction | Doc-String |
| 40 |
| Apr 21, 2026 |
| Indirect Object Identification | IOI | 40 | Apr 21, 2026 |
|---|
| Data Size Saving | Empirical Setting I | 40 | Apr 20, 2026 |
|---|
| Language Modeling | WikiText-2 | 40 | Apr 21, 2026 |
|---|
| HDL generation | RTLLM 2.0 | 40 | Apr 20, 2026 |
|---|
| HDL generation | VerilogEval 2.0 | 40 | Apr 20, 2026 |
|---|
| Packing with Traveling Thief Problem | PWT with deterministic constraints | 40 | Apr 16, 2026 |
|---|
| Landing Zone Selection | Custom Urban Delivery Dataset (test) | 40 | Apr 16, 2026 |
|---|
| Attribution Alignment | Curated Attribution Dataset (NarrativeQA + SciQ) | 40 | Apr 16, 2026 |
|---|
| Attribution faithfulness | LongRA | 40 | Apr 16, 2026 |
|---|
| Coding | HumanEval (test) | 40 | Jun 9, 2026 |
|---|
| Face Forgery Detection | P2 Hybrid, FR, FS, EFS v1 (test) | 40 | Apr 15, 2026 |
|---|
| Face Forgery Detection | P1 (FF++, DFDCP, DFD, CDF2) v1 (test) | 40 | Apr 15, 2026 |
|---|
| Closed-loop planning | NAVSIM | 40 | May 20, 2026 |
|---|
| Refusal Evaluation | OR-Bench | 40 | Apr 15, 2026 |
|---|
| Refusal Evaluation | Past Tense | 40 | Apr 15, 2026 |
|---|
| Commonsense Reasoning | Commonsense170k (test) | 40 | Jul 7, 2026 |
|---|
| Visual Reasoning | MathVerse | 40 | May 1, 2026 |
|---|
| Image Captioning | IIW-400 | 40 | Apr 15, 2026 |
|---|
| Prompt Injection Detection | PopUp attack (top visited websites) | 40 | May 15, 2026 |
|---|