Loading the SOTA2 catalog…
SOTA2 Research · benchmarks
Compare state-of-the-art methods across the tasks and datasets used to measure AI progress.
| Task | Dataset | Domain | Trend | Results | Last Update |
|---|---|---|---|---|---|
| Vulnerability Detection | PreciseBugs | 35 | Apr 23, 2026 | ||
| Social Interaction | SOTOPIA Hard | 35 | May 18, 2026 | ||
| Safety classification | WildGuard (test) | 35 | Jun 26, 2026 | ||
| Mathematical Reasoning Process Evaluation | ProcessBench (test) | 35 | May 20, 2026 | ||
| Point Cloud Registration | C3DM | 35 | Apr 21, 2026 |
| Tool Use | RoTBench Multi-turn | 35 | Apr 21, 2026 |
|---|
| Tool Use | RoTBench Single-turn | 35 | Apr 21, 2026 |
|---|
| Mathematical Reasoning | AIME | 35 | Apr 23, 2026 |
|---|
| Few-shot Text Classification | Aggregated Text Datasets | 35 | Apr 20, 2026 |
|---|
| Recommendation | Gowalla | 35 | Apr 21, 2026 |
|---|
| Personality performance evaluation | PERSONALITYBENCH | 35 | Jun 11, 2026 |
|---|
| Image Classification | CIFAR-10 (test) | 35 | Apr 16, 2026 |
|---|
| LLM Fingerprinting | LLM Lineage Verification Dataset LLaMA and Qwen-style families | 35 | Apr 15, 2026 |
|---|
| near-OOD detection | CIFAR-10, CIFAR-100, TinyImageNet Average | 35 | Apr 15, 2026 |
|---|
| far-OOD detection | Average (CIFAR-10, CIFAR-100, TinyImageNet) | 35 | Apr 15, 2026 |
|---|
| Coding | HumanEval, MBPP | 35 | Jun 1, 2026 |
|---|
| Mathematical Reasoning | AMC23 | 35 | May 8, 2026 |
|---|
| Mathematical Reasoning | Olympiad Bench | 35 | May 22, 2026 |
|---|
| Mammography Classification | VinDr | 35 | Apr 23, 2026 |
|---|
| Retinal Image Alignment | FLORI21 | 35 | Apr 14, 2026 |
|---|
| Retinal Image Alignment | KBSMC | 35 | Apr 14, 2026 |
|---|
| Text-to-SQL | Text-to-SQL Multi-sharded | 35 | Apr 10, 2026 |
|---|
| Multi-objective Recommendation | MovieLens Individual User Instances | 35 | Apr 10, 2026 |
|---|
| Pareto frontier quality evaluation | ModCloth | 35 | Apr 10, 2026 |
|---|
| Multi-objective Recommendation | ModCloth | 35 | Apr 10, 2026 |
|---|