Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Knowledge Unlearning | ZsRE | 42 | May 20, 2026 | ||
| Knowledge Unlearning | MCF | 42 | May 20, 2026 | ||
| Protein Fold Classification | TEDBench (test) | 42 | May 19, 2026 | ||
| Offline Reinforcement Learning | D4RL MuJoCo hopper-medium-replay | 42 | Jun 15, 2026 | ||
| Image Classification | ImageNet-1k | 42 | May 19, 2026 | ||
| Detoxification |
|---|
| Detoxification dataset |
| 42 |
| May 19, 2026 |
| Visual Question Answering (VQA) | MLLMU-Bench 5% (forget) | 42 | May 18, 2026 |
|---|
| Fine-grained Image Classification | 8 Fine-grained Dataset Suite Average | 42 | Jun 23, 2026 |
|---|
| Deep Research | ResearchQA | 42 | Jun 15, 2026 |
|---|
| Profile Cosine Similarity Analysis | AphasiaBank 19-symptom inventory | 42 | May 18, 2026 |
|---|
| Tool-calling | When2Call | 42 | May 15, 2026 |
|---|
| Chinese Baby Naming | CNames (test) | 42 | May 15, 2026 |
|---|
| Question Answering | BoolQ -> SQuAD -> AdversarialQA (test) | 42 | May 14, 2026 |
|---|
| Intent Detection | All Posts and Comment Mean | 42 | May 14, 2026 |
|---|
| Puzzle Solving | Sudoku | 42 | Jun 9, 2026 |
|---|
| Vulnerable Agent Identification | Vicsek environment | 42 | May 13, 2026 |
|---|
| Vulnerable Agent Identification | Battle environment | 42 | May 13, 2026 |
|---|
| Refusal Rate Evaluation | FBA (test) | 42 | May 13, 2026 |
|---|
| Conditional generation | OpenWebText | 42 | May 12, 2026 |
|---|
| Operating System GUI Agentic Reasoning | OSWorld | 42 | May 12, 2026 |
|---|
| Scheduling | FORGE-BENCH | 42 | May 12, 2026 |
|---|
| Multimodal Understanding | MMbench | 42 | May 12, 2026 |
|---|
| Jailbreak Attack | JailbreakBench (JBB) (test) | 42 | May 12, 2026 |
|---|
| Jailbreak Attack | HarmBench-191 (dev) | 42 | May 12, 2026 |
|---|
| Synthesis Planning | USPTO-190 | 42 | May 12, 2026 |
|---|