Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Referring Segmentation | RefCOCOg (test) | 52 | Jul 8, 2026 | ||
| Material Classification | FMD (test) | 52 | Mar 19, 2026 | ||
| Malicious Agent | PoisonRAG | 52 | Apr 28, 2026 | ||
| Prompt Injection | GSM8k | 52 | Apr 28, 2026 | ||
| Prompt Injection | CSQA |
| 52 |
| Apr 28, 2026 |
| Reward Modeling | RM Bench Code | 52 | Feb 26, 2026 |
|---|
| Reward Modeling | Reward Bench Math | 52 | Feb 26, 2026 |
|---|
| Medical Visual Question Answering | MedXpertQA | 52 | Jun 11, 2026 |
|---|
| Single-hop QA | NQ (Natural Questions) | 52 | May 13, 2026 |
|---|
| Linguistic and Cultural Competency | Polish Linguistic and Cultural Competency Benchmark (PLCC) | 52 | Feb 26, 2026 |
|---|
| Membership Inference Attack | MIMIR Github | 52 | Apr 1, 2026 |
|---|
| Mathematical Reasoning | Competition-level Math Benchmarks AIME24, AIME25, AMC23, MATH500, Olympiad, Minerva | 52 | Jun 26, 2026 |
|---|
| Question Answering | MultifieldQA | 52 | Feb 26, 2026 |
|---|
| Agentic Web Browsing | BrowseComp-ZH | 52 | May 19, 2026 |
|---|
| Spatial Reasoning | RealWorldQA | 52 | May 26, 2026 |
|---|
| Image Generation | GenEval (test) | 52 | Jun 25, 2026 |
|---|
| Mathematical Problem Solving | AIME | 52 | Mar 5, 2026 |
|---|
| Traffic Signal Control | Jinan-2 | 52 | Apr 29, 2026 |
|---|
| Planning | Traffic Norms Domain | 52 | Feb 26, 2026 |
|---|
| Harmful score evaluation | BeaverTails (test) | 52 | May 15, 2026 |
|---|
| GUI Grounding | OSWorld-G (test) | 52 | Feb 26, 2026 |
|---|
| Jailbreak Attack | RedTeam 2K | 52 | Mar 19, 2026 |
|---|
| Claim Verification | PerplexityAI (test) | 52 | Feb 26, 2026 |
|---|
| Retrieval-Augmented Generation (RAG) | TriviaQA | 52 | Feb 26, 2026 |
|---|
| Retrieval-Augmented Generation (RAG) | NQ | 52 | Feb 26, 2026 |
|---|