Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Machine Translation | NTREX (en->it) 128 (test) | 35 | Feb 26, 2026 | ||
| Multi-turn tool-use interaction | tau-bench | 35 | Feb 26, 2026 | ||
| Document Retrieval | COIR | 35 | Feb 26, 2026 | ||
| Multi-Hop Question Answering | Multi-Hop QA | 35 | Jun 30, 2026 | ||
| Legal Reasoning | LexEval | 35 | Apr 21, 2026 | ||
| Adversarial and Jailbreaking Attack Detection | XSTest |
| 35 |
| May 6, 2026 |
| Deep Search | BrowseComp-ZH | 35 | May 29, 2026 |
|---|
| AI-generated text detection | Essay | 35 | May 18, 2026 |
|---|
| Paraphrase Identification | PAWS | 35 | Jun 26, 2026 |
|---|
| Online Bin Packing Problem | BPP online N=1k, W=100 | 35 | Apr 8, 2026 |
|---|
| Multimodal Retrieval-Augmented Generation | MRAG | 35 | May 19, 2026 |
|---|
| Reasoning | MMLU | 35 | Mar 4, 2026 |
|---|
| Demand Forecasting | E-commerce Demand Country 01, Event-Driven Periods | 35 | Feb 26, 2026 |
|---|
| Recommendation | Yelp | 35 | Mar 26, 2026 |
|---|
| Multi-label content safety classification | Beavertails | 35 | Feb 26, 2026 |
|---|
| Code Generation | DS-1000 | 35 | Apr 23, 2026 |
|---|
| Perceptual Image Restoration | Average across datasets (combined) | 35 | Apr 8, 2026 |
|---|
| Concept Learning | Text Domains Aggregated: Academic Papers, Movie Plots, News Articles, Song Lyrics (test) | 35 | Feb 26, 2026 |
|---|
| Novel View Synthesis | DTU 1 (test) | 35 | Mar 20, 2026 |
|---|
| Common-sense Reasoning | 5 common-sense reasoning tasks Llama-3-8B | 35 | Jul 2, 2026 |
|---|
| Backdoor Attack | FaceForensics++ | 35 | Feb 26, 2026 |
|---|
| Facial Expression Recognition | AffectNet-8 | 35 | May 20, 2026 |
|---|
| Homography Estimation | GoogleMap | 35 | Mar 5, 2026 |
|---|
| Self-reenactment | HDTF | 35 | May 26, 2026 |
|---|
| Keypoint Detection | MP-100 (Split 1) | 35 | May 14, 2026 |
|---|